ETExamTower
Q9Operational Efficiency and Optimization for GenAI ApplicationsMultiple answers

A company uses Amazon Bedrock to create technical content for customers. The company has recently seen a surge in hallucinated outputs when its model produces summaries of lengthy technical documents. The outputs contain incorrect or invented details. The current solution uses a large foundation model (FM) with a basic one-shot prompt that supplies the complete document in one input. The company needs a solution that reduces hallucinations and satisfies factual-accuracy objectives. The solution must process more than 1,000 documents per hour and provide summaries within 3 seconds for each document. Which combination of solutions meets these requirements? (Choose two.)

Select 2 answers.
← → navigate · a answer
Community votes
B
47% (9)
A
42% (8)
C
11% (2)
D
0% (0)
E
0% (0)
Discussion · 9
A, B 3
The correct answers are A and B. A is correct because zero-shot chain-of-thought (CoT) instructions force the model to reason step-by-step and verify facts before generating the final summary. This reduces hallucinations by guiding the model to check each piece of information against the input content. B is correct because Retrieval Augmented Generation (RAG) with an Amazon Bedrock knowledge base allows the model to generate summaries grounded in source content. Semantic chunking and tuned embeddings ensure the model references accurate portions of the document, improving factual accuracy while enabling processing of long documents efficiently.
A, B 3
A and B seems to be the correct answer
A, B 2
Agree choose A, B
A, B 2
Correct Answer
A, B 2
✅ Reduce alucinaciones: Razonamiento paso a paso mejora precisión factual ✅ Verificación explícita: El modelo verifica hechos antes de generar ✅ Cumple latencia: Zero-shot CoT añade tokens pero mantiene <3s ✅ Escalable: Procesa 1,000+ docs/hora sin infraestructura adicional ✅ Grounding en fuente: Ancla respuestas en contenido real del documento ✅ Semantic chunking: Maneja documentos largos eficientemente ✅ Reduce alucinaciones: RAG proporciona contexto verificable ✅ Cumple latencia: Knowledge Bases optimizado para respuestas rápidas
C 1
The answer is C.
A, B 1
The combination of CoT (reasoning discipline) + RAG with chunking (grounding in source) attacks hallucinations from two complementary angles — how the model thinks and what content it works from — without sacrificing the throughput or latency requirements.
A, B 1
Hallucinations in long-document summarization are best solved by: RAG (grounding) + Structured reasoning (verification)
B, C 1
A - most probably will add more time and exceed 3 seconds D - increasing temperature will increase hallucination E - this is the current issue they try to solve