ETExamTower
Q8Operational Efficiency and Optimization for GenAI Applications

A GenAI developer is developing a Retrieval Augmented Generation (RAG)-based customer-support application that uses Amazon Bedrock foundation models (FMs). The application must process 50 GB of historical customer conversations stored as JSON files in an Amazon S3 bucket. It must use the processed data as its retrieval corpus. The application's data-processing workflow must extract relevant data from customer-support documents, remove customer personally identifiable information (PII), and generate embeddings for vector storage. The workflow must be cost-effective and complete within 4 hours. Which solution meets these requirements with the **LEAST** operational overhead?

← → navigate · a answer
Community votes
D
71% (5)
A
14% (1)
B
14% (1)
C
0% (0)
Discussion · 6
D 4
least effort and within 4 hours = D
D 3
Step Functions' Distributed Map state allows to automatically scan millions of objects in an S3 bucket and run up to 10,000 parallel workflows. This makes it easy to process 50 GB of data within a 4-hour time limit. ## Why Other Options Are Inappropriate A (Lambda Only): Managing 50 GB of data with a single or simple parallel Lambda execution would require you to code all the timeout (15 minutes) handling and API rate limiting yourself, resulting in significant operational overhead.
D 3
A has much more overhead
B 2
ChatGTP goes for B but I like D I do not know what is the good answer
D 1
Answer is D
A 1
Correct Answer