ETExamTower
Q5Data Preparation for Machine Learning (ML)

HOTSPOT - A company stores historical data in .csv files in Amazon S3. Only some of the rows and columns in the .csv files are populated. The columns are not labeled. An ML engineer needs to prepare and store the data so that the company can use the data to train ML models. Select and order the correct steps from the following list to perform this task. Each step should be selected one time or not at all. (Select and order three.) • Create an Amazon SageMaker batch transform job for data cleaning and feature engineering. • Store the resulting data back in Amazon S3. • Use Amazon Athena to infer the schemas and available columns. • Use AWS Glue crawlers to infer the schemas and available columns. • Use AWS Glue DataBrew for data cleaning and feature engineering.

Question exhibit
← → navigate · a answer
Discussion · 4
11
Order of steps: Use AWS Glue crawlers to infer schemas and available columns. Use AWS Glue DataBrew for data cleaning and feature engineering. Store the resulting data back in Amazon S3.
4
Infer the existing schema and create data catalog using Glue crawlers, than process the data with DataBrew and finally save the results in s3
1
Data brew usually comes before using glue crawlers to populate schemas / data catalogs... Order of steps: DataBrew -> store s3 -> Glue Crawler
1
Firstly, Use AWS Glue to define data schemas/tables Use Glue Data brew to perform cleaning and feature engineering Upload the transformed data back into s3