Q88Design Secure Architectures
A hospital recently implemented a RESTful API by using Amazon API Gateway and AWS Lambda. The hospital uses API Gateway and Lambda to upload reports in PDF format and JPEG format. The hospital must update the Lambda code to identify protected health information (PHI) in the reports. Which solution will meet these requirements with the **LEAST operational overhead**?
← → navigate · a answer
Community votes
Discussion · 20
C 21
The correct solution is C: Use Amazon Textract to extract the text from the reports. Use Amazon Comprehend Medical to identify the PHI from the extracted text.
Option C: Using Amazon Textract to extract the text from the reports, and Amazon Comprehend Medical to identify the PHI from the extracted text, would be the most efficient solution as it would involve the least operational overhead. Textract is specifically designed for extracting text from documents, and Comprehend Medical is a fully managed service that can accurately identify PHI in medical text. This solution would require minimal maintenance and would not incur any additional costs beyond the usage fees for Textract and Comprehend Medical.
8
Option A: Using existing Python libraries to extract the text and identify the PHI from the text would require the hospital to maintain and update the libraries as needed. This would involve operational overhead in terms of keeping the libraries up to date and debugging any issues that may arise.
Option B: Using Amazon SageMaker to identify the PHI from the extracted text would involve additional operational overhead in terms of setting up and maintaining a SageMaker model, as well as potentially incurring additional costs for using SageMaker.
Option D: Using Amazon Rekognition to extract the text from the reports would not be an effective solution, as Rekognition is primarily designed for image recognition and would not be able to accurately extract text from PDF or JPEG files.
C 4
Both Rekognition and Textract possess the ability to detect text within images, yet they are optimized for differing applications.
Rekognition specializes in identifying text located spatially within an image, for instance, words displayed on street signs, t-shirts, or license plates. Its typical use cases encompass visual search, content filtering, deriving insights from content, among others. However, it's not the ideal choice for images containing more than 100 words, as this exceeds its limitation.
On the other hand, Textract is tailored more towards processing documents and PDFs, offering a comprehensive suite for Optical Character Recognition (OCR). It proves useful in scenarios involving financial reports, medical records, receipts, ID documents, and more.
3
D is wrong only because Amazon Rekognition doesn't read text, only explicit image contents.
C 3
Textract = Extract text from PDF/iamges
Comprehend Medical = PHI
ABD are wrong products for this requirement so won't achieve the results
3
Selected Answer: C
Amazon Textract is a machine learning (ML) service that automatically extracts text, handwriting, and data from scanned documents.
C 2
• Amazon Textract: This program is made to extract text and data from scanned documents, such as pictures and PDFs. It helps to retain the formatting of the report by automatically extracting text while preserving the document's layout.
Identifying and extracting medical information, including protected health information (PHI), from unstructured text is the specialty of Amazon Comprehend Medical. Medical entities that are frequently included in reporting on healthcare, such as ailments, drugs, and more, can be recognized by it.
2
The correct solution is C: Use Amazon Textract to extract the text from the reports. Use Amazon Comprehend Medical to identify the PHI from the extracted text.
Option C: Using Amazon Textract to extract the text from the reports, and Amazon Comprehend Medical to identify the PHI from the extracted text, would be the most efficient solution as it would involve the least operational overhead. Textract is specifically designed for extracting text from documents, and Comprehend Medical is a fully managed service that can accurately identify PHI in medical text. This solution would require minimal maintenance and would not incur any additional costs beyond the usage fees for Textract and Comprehend Medical.
C 2
Use Textract to extract text from medical reports (PDFs, JPEGs) and Comprehend Medical to detect PHI — with minimal code and maximum accuracy.
2
with the choices here, I would go with C, but if offered, I would use amazon textract for the text and use Macie to do the scanning of text files, not comprehend.
C 2
Ans C - Textract to 'read' data; Comprehend to assess whether its PHI
2
Option C is the right answer.
C 2
C leverages capabilities of Textract, which is a service that automatically extracts text and data from documents, including PDF and JPEG. By using Textract, hospital can extract text content from reports without need for additional custom code or libraries.
Once text is extracted, hospital can then use Comprehend Medical, a natural language processing service specifically designed for medical text, to analyze and identify PHI. It can recognize medical entities such as medical conditions, treatments, and patient information.
A. suggests using existing Python libraries, which would require hospital to develop and maintain custom code for text extraction and PHI identification.
B and D involve using Textract along with SageMaker or Rekognition, respectively, for PHI identification. While these options could work, they introduce additional complexity by incorporating machine learning models and training.
C 1
Option C
1
Key word: hospital!
C 1
Textract is more suitable than Rekognition as it is build to scan text documents and Comprehend helps identify PHI.
C 1
Here's why:
Amazon Textract has built-in support to extract text from PDFs and images, eliminating the need to build this yourself with Python libraries.
Amazon Comprehend Medical has pre-trained machine learning models to identify PHI entities out-of-the-box, avoiding the need to train your own SageMaker model.
Using these fully managed AWS services minimizes operational overhead of maintaining machine learning models yourself.
1
WHY OPTION D IS WRONG
1
B/C you use TextTract to extract text not Rekognition.
1
Answer C: