Q10Data Preparation for Machine Learning (ML)
Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. Which AWS service or feature can aggregate the data from the various data sources?
← → navigate · a answer
Community votes
Discussion · 19
A 17
Selected Answer: A
Amazon EMR with Spark is a strong choice for aggregating, processing, and transforming large datasets from multiple sources (e.g., Amazon S3 and on-premises MySQL database). Spark jobs can handle both structured and unstructured. While Lake Formation is great for managing data lakes, it doesn’t provide the ETL and data processing capabilities needed to aggregate and transform datasets from multiple sources.
D 10
Selected Answer: D
Is it D? AWS Lake Formation ? EMR Spark jobs is more manual.
D 7
Selected Answer: D
AWS Lake Formation is a service built to aggregate, catalog, and manage data from multiple data sources, including on-premises databases and Amazon S3, making it an ideal choice for this scenario.
While Amazon EMR with Apache Spark is powerful for processing and analyzing large datasets, it focuses more on data processing than on data aggregation and cataloging. It doesn't inherently manage interdependencies or schema enforcement
D 6
Selected Answer: D
Lake Formation is the correct answer
D 6
Selected Answer: D
The answer is D (Lake Formation). This is because EMR Spark does not natively support pulling data from on-prem DB as its data source. You would need DataSync or something else for that. On the other hand, Lake Formation fulfills all use cases documented clearly as links shown below.
https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html#lake-formation-features
https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-plan-get-data-in.html
D 5
Selected Answer: D
Data Lake is used for data discovery
D 4
Selected Answer: D
The correct answer is D. AWS Lake Formation.
Explanation:
AWS Lake Formation is made for aggregating, organizing, and securing large datasets from multiple sources (e.g., S3, on-premises databases). It simplifies the creation of a centralized data lake, making it easier to integrate and analyze diverse data formats. This is especially useful for tasks like fraud detection, where data comes from different sources.
Amazon EMR Spark jobs (Option A) is more suited for large-scale data processing and analytics. While it can process and transform data, it requires more operational effort to configure and manage compared to AWS Lake Formation.
Why AWS Lake Formation?
Aggregates and organizes data from S3 and MySQL easily.
Offers integrated data cataloging for better feature engineering.
Reduces operational overhead compared to setting up EMR.
D 4
Selected Answer: D
Lake Formation ain't only the storage service - it actually an umbrella over most AWS Glue Offerings. And cause it also provides fully serverless Spark, it seems to be better option than EMR. This point is quite tricky as often when someone refers to LF, they mean the governance part only.
A 3
Selected Answer: A
A. Amazon EMR Spark jobs
A 2
Selected Answer: A
Amazon EMR is correct answer for aggregation
A 2
Selected Answer: A
EMR with Spark can aggregate large datasets from multiple sources, including S3 and on-premises MySQL.
A 1
Selected Answer: A
following strictly by question "can aggregate data" - it's A, Spark , indisputable. DataLake - it's the Storage, repository, something static conception and an idea. I can not understand the votes for D
1
really spark can not? What about the : Outposts arch, Direct Connection?
D 1
Selected Answer: D
It allows the ML engineer to bring together transaction logs, customer profiles, and MySQL data into one governed data lake (usually in S3) where the dataset can be queried, joined, and used for training the fraud detection model.
A 1
Selected Answer: A
Lake Formation does not support aggregation, transformation, joins, feature engineering.
EMR with Spark is right choice
D 1
Selected Answer: D
AWS Lake Formation is built specifically to build, manage, and aggregate data from multiple sources into a centralized data lake on Amazon S3
A 1
Selected Answer: A
EMR because it can do all the heavy lifting where-as AWS LakeFormation is more targetted towards permission management for data lakes.
1
I think it is D, because EMR with SPARK is more of a compute engine, not really a central aggregation service.
A 1
Selected Answer: A
EMR ... Lake Formation isnt handling the data transformations.