ETExamTower
Q21Data Ingestion and Transformation

A company is migrating on-premises workloads to AWS. The company wants to reduce overall operational overhead. The company also wants to explore serverless options. The company's current workloads use Apache Pig, Apache Oozie, Apache Spark, Apache Hbase, and Apache Flink. The on-premises workloads process petabytes of data in seconds. The company must maintain similar or better performance after the migration to AWS. Which extract, transform, and load (ETL) service will meet these requirements?

← → navigate · a answer
Community votes
B
69% (9)
A
31% (4)
C
0% (0)
D
0% (0)
Discussion · 17
B 19
Selected Answer: B Glue is like the better-looking but weaker sibling of EMR. So when it comes to petabyte-scale workloads, let EMR handle the job and keep Glue away from the action.
A 3
Selected Answer: A Glue is Serverless :)
B 3
Selected Answer: B EMR offers a managed Hadoop framework that natively supports Apache Pig, Oozie, Spark, and Flink. This lets the company move its existing workloads with minimal code changes, which cuts development effort
2
A. AWS Glue: AWS Glue is a fully managed extract, transform, and load (ETL) service offered by Amazon Web Services (AWS). It lets users prepare and load data for analytics purposes B. Amazon EMR: Amazon Elastic MapReduce (EMR) is a cloud-based big data platform offered by AWS. It lets users process and analyze large amounts of data with popular frameworks such as Apache Hadoop, Apache Spark, Apache Hive, Apache HBase, and more. https://docs.aws.amazon.com/emr/index.html https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-best-practices.html https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-manage.html https://docs.aws.amazon.com/emr/latest/DeveloperGuide/emr-developer-guide.html As per the AWS/Amazon docs, option B specifically calls out the exact features/options that the question asked about directly.
2
- While AWS Glue is a fully managed ETL service and does offer serverless capabilities, it might not deliver the same level of performance and flexibility as Amazon EMR for petabyte-scale workloads with complex processing needs. - AWS Glue is built for data integration, cataloging, and ETL jobs but may not be as good for heavy-duty processing tasks that need frameworks like Apache Spark, Apache Flink, etc., which are commonly used for large-scale data processing. - Documentation on AWS Glue can be found in the AWS Glue Developer Guide https://docs.aws.amazon.com/glue/index.html.
B 2
Selected Answer: B https://docs.aws.amazon.com/ja_jp/emr/latest/ManagementGuide/emr-what-is-emr.html
B 2
Selected Answer: B That's exactly what EMR is for. "Amazon EMR is the industry-leading cloud big data solution for petabyte-scale data processing, interactive analytics, and machine learning using open-source frameworks such as Apache Spark, Apache Hive, and Presto." https://aws.amazon.com/emr/
A 2
Selected Answer: A I think it is A, Glue • Amazon EMR is used for petabyte-scale data collection and data processing. • AWS Glue is used as a serverless and managed ETL service, and also used for managing data quality with AWS Glue Data Quality.
2
Discarded, not 'discarted'. 'Discarted' isn't a word.
B 2
Selected Answer: B Amazon EMR Serverless is a deployment option for Amazon EMR that provides a serverless runtime environment. This simplifies the operation of analytics applications that use the latest open-source frameworks, such as Apache Spark and Apache Hive. With EMR Serverless, you don’t have to configure, optimize, secure, or operate clusters to run applications with these frameworks. https://docs.aws.amazon.com/emr/latest/EMR-Serverless-UserGuide/emr-serverless.html
A 1
Selected Answer: A Serverless: AWS Glue is a fully managed, serverless ETL service that automates data discovery, preparation, and transformation, helping reduce operational overhead.Integration with Big Data Tools: It integrates well with different AWS services and supports Spark jobs for ETL purposes, which fits Apache Spark workloads well.Performance: AWS Glue can handle large-scale ETL workloads, and it is built to manage petabytes of data efficiently, comparable to the performance of on-premises solutions.While B. Amazon EMR could also be considered for its flexibility in handling big data workloads using tools like Apache Spark, it requires more management and does not match the serverless requirement as closely as AWS Glue. Therefore, AWS Glue is the most suitable choice given the constraints and requirements.
1
The company also wants to explore serverless options. ? Glue (A). or EMR Serverless
A 1
Selected Answer: A Glue. It mentions "serverless" so EMR is discarted. The mention of Spark, Hbase, etc is there to confuse you, because it doesn't say they need to keep using them. Glue can run Spark with "glueContext" (similar to SparkContext) for reading tables, files and creating frames.
1
B. Amazon EMR Serverless is a deployment option for Amazon EMR that provides a serverless runtime environment. This simplifies the operation of analytics applications that use the latest open-source frameworks, such as Apache Spark and Apache Hive. With EMR Serverless, you don’t have to configure, optimize, secure, or operate clusters to run applications with these frameworks.
B 1
Selected Answer: B Apache = EMR
B 1
Selected Answer: B Glue doesnt natively support Pig, HBase and Flink.
B 1
Selected Answer: B EMR for big data.