Q35Design High-Performing Architectures
A solutions architect manages an analytics application. The application stores large volumes of semistructured data in an Amazon S3 bucket. The solutions architect wants to use parallel data processing so the data can be processed more quickly. The solutions architect also wants to use information stored in an Amazon Redshift database to enrich the data. Which solution meets these requirements?
← → navigate · a answer
Community votes
Discussion · 21
B 18
Option B is the correct solution that meets the requirements:
Use Amazon EMR to process the semi-structured data in Amazon S3. EMR provides a managed Hadoop framework optimized for processing large datasets in S3.
EMR supports parallel data processing across multiple nodes to speed up the processing.
EMR can integrate directly with Amazon Redshift using the EMR-Redshift integration. This allows querying the Redshift data from EMR and joining it with the S3 data.
This enables enriching the semi-structured S3 data with the information stored in Redshift
9
By combining AWS Glue and Amazon Redshift, you can process the semistructured data in parallel using Glue ETL jobs and then store the processed and enriched data in a structured format in Amazon Redshift. This approach allows you to perform complex analytics efficiently and at scale.
B 8
A has a pitfall, "use Amazon Athena to PROCESS the data". With Athena you can query, not process, data.
C is wrong because Kinesis has no place here.
D is wrong because it does not process the Redshift data, and Glue does ETL, not analyze
Thus it's B. EMR can use semi-structured data from from S3 and structured data from Redshift and is ideal for "parallel data processing" of "large amounts" of data.
A 4
semi-structure supported by Athena not by EMR
B 4
D: not relevant, data is semistructured and Glue is more batch than stream data
A: not correct, Athena is for querying data
B & C look ok but C is out => redundant with Kinesis data stream; EMR already processed data as input into Redshift for parallel processing
Only B is most logical
B 3
Key requirement: parallel data processing
parallel data processing is EMR (Kind of Apache Hadoop) so it only leave B and C
C is Kinesis to Redshift which is pointless logic here
B EMR for S3 and EMR for Redshift gives maximum parallel processing here
3
Athena is not designed for parallel data processing. So it's B
B 3
large amount of data + parallel data processing = EMR
B 2
From this documentation looks like EMR cannot interface with S3.
https://aws.amazon.com/emr/
I will settle with option A.
B 2
Choose option B.
Option A is not correct. Amazon Athena is suitable for querying data directly from S3 using SQL and allows parallel processing of S3 data.
AWS Glue can be used for data preparation and enrichment but might not directly integrate with Amazon Redshift for enrichment.
B 2
For those answering A, AWS Glue can directly query S3, it can't use Athena as a source of data. The questions say the Redshift data should be user to "enrich" which means thats the redshift data needs to be "added" to the s3 data. A doesn't allow that.
2
"Hadoop helps you turn petabytes of un-structured or semi-structured data into useful insights about your applications or users."
https://aws.amazon.com/emr/features/hadoop/?nc1=h_ls
2
Y, but A says "process", not "query" data with Athena.
2
"Hadoop [as used by EMR] helps you turn petabytes of un-structured or semi-structured data into useful insights"
https://aws.amazon.com/emr/features/hadoop/
2
Of course EMR can access S3
https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-plan-file-systems.html
A 1
Athena and Redshift both do SQL query
A 1
Answer is A
1
i think a is correct
semistructured data ==> Athena
A 1
athena for s3
A 1
Amazon Athena is an interactive query service that makes it easy to analyze data directly in Amazon Simple Storage Service (Amazon S3) using standard SQL.
1
Selected Answer: D
Glue use apache pyspark cluster for parallel processing. EMR or Glue are possible options. Glue is serverless so better use this plus pyspark is in memory parallel processing.