ETExamTower
Q26Data Store Management

A company maintains an Amazon Redshift provisioned cluster that the company uses for extract, transform, and load (ETL) operations to support critical analysis tasks. A sales team within the company maintains a Redshift cluster that the sales team uses for business intelligence (BI) tasks. The sales team recently requested access to the data that is in the ETL Redshift cluster so the team can perform weekly summary analysis tasks. The sales team needs to join data from the ETL cluster with data that is in the sales team's BI cluster. The company needs a solution that will share the ETL cluster data with the sales team without interrupting the critical analysis tasks. The solution must minimize usage of the computing resources of the ETL cluster. Which solution will meet these requirements?

← → navigate · a answer
Community votes
A
67% (6)
D
33% (3)
B
0% (0)
C
0% (0)
Discussion · 14
7
Sorry, I meant to choose A but selected D
A 6
Selected Answer: A A: redshift data sharing: https://docs.aws.amazon.com/redshift/latest/dg/data_sharing_intro.html With data sharing, you can securely and easily share live data across Amazon Redshift clusters. B: materialized view is only within 1 redshift cluster, across different tables
A 5
Selected Answer: A At first I would pick B, but that would definitely use more resources.
D 5
Selected Answer: D In my view, using Redshift Data Sharing will use fewer resources. 'D' envolves using a S3 bucket.
A 4
Selected Answer: A To share data between Redshift clusters and meet the requirements of sharing ETL cluster data with the sales team without interrupting critical analysis tasks and minimizing the usage of the ETL cluster's computing resources, Redshift Data Sharing is the way to go https://docs.aws.amazon.com/redshift/latest/dg/data_sharing_intro.html "Supporting different kinds of business-critical workloads – Use a central extract, transform, and load (ETL) cluster that shares data with multiple business intelligence (BI) or analytic clusters. This approach provides read workload isolation and chargeback for individual workloads. You can size and scale your individual workload compute according to the workload-specific requirements of price and performance"
A 4
Selected Answer: A Typical datasharing use case in Redshift. The question says - 'team needs to join data from the ETL cluster with data that is in the sales team's BI cluster.' This is possible with datashare.
D 3
Selected Answer: D The spectrum table is accessed from the sales cluster with zero impact on the ETL cluster.
A 3
Selected Answer: A Typical Redshift data sharing use case
2
Overall, although both options provide ways to share data between the ETL and BI clusters, Option D gives a more robust and scalable solution that minimizes the impact on the ETL cluster's resources and offers greater flexibility and independence for the sales team's analysis tasks. By unloading a copy of the data from the ETL cluster to Amazon S3 and using Redshift Spectrum for querying, the solution follows AWS best practices for managing data and resource usage in Amazon Redshift clusters. It makes sure that critical analysis tasks are not interrupted while still giving the sales team the needed access to perform their analysis tasks efficiently.
2
key words: "weekly" "The solution must minimize usage of the computing resources of the ETL cluster." Answer:D
2
It seems the performance of the critical ETL cluster should not be impacted when using data sharing, so the answer is probably A: https://docs.aws.amazon.com/redshift/latest/dg/data_sharing_intro.html Supporting different kinds of business-critical workloads – Use a central extract, transform, and load (ETL) cluster that shares data with multiple business intelligence (BI) or analytic clusters. This approach provides read workload isolation and chargeback for individual workloads. You can size and scale your individual workload compute according to the workload-specific requirements of price and performance. https://docs.aws.amazon.com/redshift/latest/dg/considerations.html The performance of queries on shared data depends on the compute capacity of the consumer clusters.
1
Options A, B, and C require giving the sales team direct access to the ETL cluster, which could possibly affect the performance of the ETL cluster and interfere with its critical analysis tasks. Option D offers a more isolated and scalable approach by using Amazon S3 and Redshift Spectrum for data sharing while minimizing the usage of the ETL cluster's computing resources. https://docs.aws.amazon.com/redshift/latest/dg/c-using-spectrum-sharing-data.html https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-design-tables.html
D 1
Selected Answer: D "The solution must minimize usage of the computing resources of the ETL cluster." That is the key point. You shouldn't use the ETL cluster, so unload the data to S3 and run the queries in a separate Redshift Spectrum database. The ETL cluster does nothing in the meantime.
1
A as Redshift data sharing lets you share live data across Redshift clusters without needing to duplicate the data. This feature lets the sales team access the data from the ETL cluster directly without interrupting the critical analysis tasks or overloading the ETL cluster's resources. The sales team can join this shared data with their own data in the BI cluster efficiently.