ETExamTower
Q27Data Ingestion and Transformation

A data engineer needs to join data from multiple sources to perform a one-time analysis job. The data is stored in Amazon DynamoDB, Amazon RDS, Amazon Redshift, and Amazon S3. Which solution will meet this requirement MOST cost-effectively?

← → navigate · a answer
Community votes
C
100% (3)
A
0% (0)
B
0% (0)
D
0% (0)
Discussion · 5
C 7
Selected Answer: C I’d pick C because Federated Query is the usual fit for this purpose. Besides, we don’t need to add/duplicate resources in S3. But I see that, because Athena is more optimized for S3, it can be seen as a tricky question, since there may be more trade-offs to think about, such as data governance that are easier if data is centralized in S3 in my opinion.
C 4
Selected Answer: C Serverless Processing: Athena is a serverless query service, so you only pay for the queries you run. This removes the need to provision and manage compute resources like in EMR clusters, making it ideal for one-time jobs. Federated Query Capability: Athena Federated Query lets you directly query data from different sources like DynamoDB, RDS, Redshift, and S3 without physically moving the data. This avoids data movement costs and simplifies the analysis process. Reduced Cost for Large Datasets: Compared to copying data to S3, which can be expensive for large datasets, Athena Federated Query skips unnecessary data movement, lowering overall costs.
2
Amazon Athena Federated Query lets you query data from multiple federated data sources including relational databases, NoSQL databases, and object stores directly from Athena. While this may seem like an efficient way to join data from different sources without having to copy data into Amazon S3, it's important to think about the cost implications. AWS documentation on Amazon Athena Federated Query [1] explains that while Federated Query lets you query external data sources without moving the data, it does not remove data transfer costs. Depending on the data sources involved (such as Amazon RDS, DynamoDB, etc.), there could be data transfer costs tied to querying data directly from these sources. [1] Amazon Athena Federated Query Documentation: https://docs.aws.amazon.com/athena/latest/ug/federated-data-sources.html
1
Point: "perform a one-time analysis job" Option C (Amazon Athena Federated Query) may look appealing, but it's usually better for querying data from external sources without copying the data into S3. However, because the data is already in AWS services, copying it to S3 and using Athena directly would likely be more cost-effective.
1
1. Data Storage Costs: Amazon S3 storage is usually cheaper than the other AWS storage options like Amazon Redshift or Amazon RDS. 2. Compute Costs: Amazon Athena is a serverless query service that lets you query data directly from S3 without needing to provision or manage infrastructure. You only pay for the queries you run, which can be more cost-effective than provisioning an EMR cluster (option A) or using Redshift Spectrum (option D), both of which require compute resources you might not fully use. 3. Data Transfer Costs: Option B means copying the data once into S3, and after that there are no extra data transfer costs when querying the data with Athena. By comparison, options A and D would have data transfer costs as data moves between different services. Amazon Athena Pricing: https://aws.amazon.com/athena/pricing/ Amazon S3 Pricing: https://aws.amazon.com/s3/pricing/