ETExamTower
Q13Data Preparation for Machine Learning (ML)

Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. Before the ML engineer trains the model, the ML engineer must resolve the issue of the imbalanced data. Which solution will meet this requirement with the LEAST operational effort?

← → navigate · a answer
Community votes
D
100% (5)
A
0% (0)
B
0% (0)
C
0% (0)
Discussion · 4
D 6
Selected Answer: D Least effort https://aws.amazon.com/blogs/machine-learning/balance-your-data-for-machine-learning-with-amazon-sagemaker-data-wrangler/
D 1
Selected Answer: D Both Glue DataBrew and Data Wrangler let you prepare data for ML with no-code/low-code (aka low ops effort). However, Data Wrangler has built-in transformations for balancing datasets (random oversampling, random undersampling and smote) https://docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler-transform.html#data-wrangler-transform-balance-data while DataBrew doesn't have a built-in recipe step for balancing datasets, actually it offers a smaller set of data science recipe steps limited to binarization, bucketization, categorical mapping, one-hot encoding, scaling, skewness and tokenization https://docs.aws.amazon.com/databrew/latest/dg/recipe-actions.data-science.html
D 1
Selected Answer: D SageMaker Data Wrangler includes a "balance data" operation built to handle class imbalance.
D 1
Selected Answer: D SageMaker Data Wrangler provides balanced data handling and deals with imbalance with minimal effort.