Q33Data Operations and SupportMultiple answers
A company is building an analytics solution. The solution uses Amazon S3 for data lake storage and Amazon Redshift for a data warehouse. The company wants to use Amazon Redshift Spectrum to query the data that is in Amazon S3. Which actions will provide the FASTEST queries? (Choose two.)
Select 2 answers.
← → navigate · a answer
Community votes
Discussion · 7
B, C 6
Selected Answer: BC
https://docs.aws.amazon.com/redshift/latest/dg/c-spectrum-external-performance.html
B, C 5
Selected Answer: BC
B. Use a columnar storage file format: This is a very good approach. Columnar storage formats like Parquet and ORC are strongly recommended for use with Redshift Spectrum. They store data in columns, which lets Spectrum scan only the columns needed for a query, greatly improving query performance and reducing the amount of data scanned.
C. Partition the data based on the most common query predicates: Partitioning data in S3 based on commonly used query predicates (like date, region, etc.) lets Redshift Spectrum skip large portions of data that are irrelevant to a specific query. This can bring major performance gains, especially for large datasets.
B, C 3
Selected Answer: BC
Redshift Spectrum is optimized for querying data stored in columnar formats like Parquet or ORC.
These formats store each data column separately, so Redshift Spectrum only needs to scan the relevant columns for a specific query, which greatly improves performance compared to row-oriented formats
Partitioning organizes data files in S3 based on specific column values (e.g., date,
region). When your queries filter or join data using these partitioning columns (common query predicates), Redshift Spectrum can quickly find the relevant data files, minimizing the amount of data scanned and speeding up query execution
1
1. **Columnar Storage File Format**:
According to AWS documentation, columnar storage file formats like Apache Parquet and Apache ORC are recommended for optimizing query performance with Amazon Redshift Spectrum. They note that these formats are highly efficient for selective column reads, which matches how analytical queries usually work. This can be found in the AWS documentation for Amazon Redshift Spectrum under "Choosing Data Formats": https://docs.aws.amazon.com/redshift/latest/dg/c-using-spectrum.html#spectrum-columnar-storage
1
2. **Partitioning**:
AWS documentation for Amazon Redshift Spectrum emphasizes the value of partitioning data based on commonly used query predicates to improve query performance. By partitioning data, Redshift Spectrum can prune unneeded partitions during query execution, reducing the amount of data scanned and improving overall query performance. This guidance is available in the AWS documentation for Amazon Redshift Spectrum under "Using Partitioning to Improve Query Performance": https://docs.aws.amazon.com/redshift/latest/dg/c-using-spectrum-partitioning.html
B, C 1
Selected Answer: BC
https://aws.amazon.com/blogs/big-data/10-best-practices-for-amazon-redshift-spectrum/
B, C 1
Selected Answer: BC
Partitioning helps filter the data and columnar storage is optimized for analytical (OLAP) queries