ETExamTower
Q32Data Ingestion and Transformation

A company is migrating a legacy application to an Amazon S3 based data lake. A data engineer reviewed data that is associated with the legacy application. The data engineer found that the legacy data contained some duplicate information. The data engineer must identify and remove duplicate information from the legacy application data. Which solution will meet these requirements with the LEAST operational overhead?

← → navigate · a answer
Community votes
B
80% (4)
A
20% (1)
C
0% (0)
D
0% (0)
Discussion · 5
B 6
Selected Answer: B Option B, using an AWS Glue ETL job with the FindMatches ML transform, is likely to satisfy the requirements with the least operational overhead. This approach uses a managed service (AWS Glue) and includes a built-in ML transform made specifically for deduplication, which reduces the need for manual setup, maintenance, and machine learning expertise.
B 4
Selected Answer: B B. https://docs.aws.amazon.com/glue/latest/dg/machine-learning.html "Find matches Finds duplicate records in the source data. You train this machine learning transform by labeling example datasets to show which rows match. The machine learning transform learns which rows should be matches the more you teach it with example labeled data."
A 3
Selected Answer: A I disagree with B. That option takes extra effort just to train the ML model with labeled data. Option A is as simple as using the robust pandas library
1
Remove duplicates from already migrated data - probably D. Remove duplicates from data before migration - A is preferable.
B 1
Selected Answer: B 100% B