Q20Data Preparation for Machine Learning (ML)
An ML engineer needs to process thousands of existing CSV objects and new CSV objects that are uploaded. The CSV objects are stored in a central Amazon S3 bucket and have the same number of columns. One of the columns is a transaction date. The ML engineer must query the data based on the transaction date. Which solution will meet these requirements with the LEAST operational overhead?
← → navigate · a answer
Community votes
Discussion · 6
A 2
Selected Answer: A
Basic CTAS usage
A 2
Selected Answer: A
Athena lets you query data stored in Amazon S3 directly using SQL without needing data movement or transformation. CTAS (CREATE TABLE AS SELECT): Creates a new table based on a filtered or transformed dataset, such as transaction dates, and stores the output in S3.
Why Not the Other Options?
B. S3 Object Lambda is meant for on-the-fly data transformation, not efficient data querying. Adding replication increases complexity without directly solving the querying requirement.
C. Glue is good for complex ETL workflows, but it adds significant operational overhead for a task that Athena can handle more simply.
D. Firehose is designed for streaming data, not processing large existing datasets.
A 1
Selected Answer: A
Using Amazon Athena with a CREATE TABLE AS SELECT (CTAS) statement is the easiest and most efficient way to query the CSV objects based on the transaction date, while needing minimal operational effort.
A 1
Selected Answer: A
A. Yes, Athena is the right service to query data in S3.
B. No, maybe this could also work, but it is pretty cumbersome
C. No, SparkSQL can be used to query files on data, but it is more work than Athena and creating a new S3 bucket is not necessary
D. No, Data Firehose cannot consume from S3 directly
A 1
Selected Answer: A
None of the answers are right. Amazon CTAS cannot create a table that points directly to S3 in a single query/step.
A 1
Selected Answer: A
Least effort is to set up an external table on top of the S3 data with Athena and query the data directly from Athena, the date column can be used to create partitions in the table