ETExamTower
Q11ML Solution Monitoring, Maintenance, and Security

Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. After the data is aggregated, the ML engineer must implement a solution to automatically detect anomalies in the data and to visualize the result. Which solution will meet these requirements?

← → navigate · a answer
Community votes
C
100% (5)
A
0% (0)
B
0% (0)
D
0% (0)
Discussion · 4
C 5
Selected Answer: C https://docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler-analyses.html "Amazon SageMaker Data Wrangler includes built-in analyses that help you generate visualizations and data analyses in a few clicks. " This question is tricky, since it makes you think you need Quicksight for the "visualization' part.
C 4
Selected Answer: C SageMaker Data Wangler identifies anomalies as part of the Data Quality and Insights Report (https://docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler-data-insights.html) and offers different options for data visualization - https://docs.aws.amazon.com/sagemaker/latest/dg/data-wrangler-analyses.html
C 2
Selected Answer: C Why Transform Categorical Data into Numerical Data? Machine learning algorithms generally need categorical data to be turned into numerical representations (e.g., one-hot encoding or embeddings) for training. Turning numerical data into categorical data is unnecessary unless the problem explicitly calls for it (e.g., binning for some specific applications). Why Use SageMaker Data Wrangler? Minimal Operational Overhead: Amazon SageMaker Data Wrangler offers a user-friendly interface to clean, preprocess, and transform data without having to write custom code. Comprehensive Data Handling: Supports data sources like S3 and on-premises databases, and can handle both categorical and numerical data transformations efficiently. Why Not AWS Glue? AWS Glue is better suited for large-scale ETL (Extract, Transform, Load) operations, such as schema inference or combining large datasets. It has higher operational overhead for specific ML data preprocessing tasks compared to SageMaker Data Wrangler.
C 1
Selected Answer: C Data Wrangler