ETExamTower
Q21Data Preparation for Machine Learning (ML)

A company has a large, unstructured dataset. The dataset includes many duplicate records across several key attributes. Which solution on AWS will detect duplicates in the dataset with the LEAST code development?

← → navigate · a answer
Community votes
D
60% (6)
A
20% (2)
C
20% (2)
B
0% (0)
Discussion · 9
D 9
Selected Answer: D AWS Glue FindMatches is built to identify duplicate or matching records in datasets without needing labeled training data. It uses machine learning to find fuzzy matches and lets you customize the matching process, which makes it a good fit for this scenario.
D 7
Selected Answer: D https://aws.amazon.com/about-aws/whats-new/2021/11/aws-glue-findmatches-new-data-existing-dataset/ "allows you to identify duplicate or matching records in your dataset"
D 4
Selected Answer: D The AWS Glue FindMatches transform is the best fit because it is specifically built to detect duplicates, needs minimal development effort, and scales well for large datasets.
A 2
Selected Answer: A I'm not sure but I think this is the right answer.
C 2
Selected Answer: C I would argue you can should use Data Wrangler if you want the least code development. See https://docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler-data-insights.html#data-wrangler-data-insights-samples "You can remove duplicate samples from the dataset using the Drop duplicates transform under Manage rows."
D 1
Selected Answer: D AWS Glue FindMatches needs the least code development because it's a purpose-built transform made specifically for matching similar records. You can train it with examples of matches and non-matches, then apply it to your whole dataset without writing complex matching algorithms. Amazon SageMaker Data Wrangler can help with data preparation, but it would require more manual coding to build custom deduplication logic.
D 1
Selected Answer: D "The FindMatches transform enables you to identify duplicate or matching records in your dataset, even when the records do not have a common unique identifier and no fields match exactly. This will not require writing any code or knowing how machine learning works. " [+] https://docs.aws.amazon.com/glue/latest/dg/machine-learning.html Why its not C. The question asked to identify, not drop. Mentioned by another user.
C 1
Selected Answer: C It says unstructured data. Glue FindMatches is only meant for structured data.
A 1
Selected Answer: A Go look this up... Amazon Mechanical Turk (MTurk) is a commonly used and effective tool for finding and removing duplicates in unstructured data. The platform uses human intelligence to handle tasks that automated systems find difficult, such as recognizing nuances in human-generated text or spotting near-identical images or listings.