Q30Data Operations and SupportMultiple answers
A company uses an Amazon QuickSight dashboard to monitor usage of one of the company's applications. The company uses AWS Glue jobs to process data for the dashboard. The company stores the data in a single Amazon S3 bucket. The company adds new data every day. A data engineer discovers that dashboard queries are becoming slower over time. The data engineer determines that the root cause of the slowing queries is long-running AWS Glue jobs. Which actions should the data engineer take to improve the performance of the AWS Glue jobs? (Choose two.)
Select 2 answers.
← → navigate · a answer
Community votes
Discussion · 6
A, B 11
Selected Answer: AB
A. Partition the data that is in the S3 bucket. Organize the data by year, month, and day.
• Partitioning data in Amazon S3 can greatly improve query performance. By arranging the data by year, month, and day, AWS Glue and Amazon QuickSight can scan only the partitions that matter, which cuts down the amount of data read and processed. This is especially useful for time-series data, where queries usually focus on specific time ranges.
B. Increase the AWS Glue instance size by scaling up the worker type.
• Scaling up the worker type can give the AWS Glue jobs more compute resources, allowing them to process data more quickly. This can be especially helpful when working with large datasets or complex transformations. It’s important to watch the performance gains and the cost impact of scaling up.
2
1. **Partition the Data in Amazon S3**:
- AWS documentation on optimizing Amazon S3 performance: https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimizing-performance.html
- AWS Glue documentation on partitioning data for AWS Glue jobs: https://docs.aws.amazon.com/glue/latest/dg/how-it-works.html#how-partitioning-works
- Best practices for partitioning in Amazon S3: https://docs.aws.amazon.com/AmazonS3/latest/userguide/best-practices-partitioning.html
2. **Optimizing AWS Glue Job Settings**:
- AWS Glue documentation on optimizing job performance: https://docs.aws.amazon.com/glue/latest/dg/best-practices.html
- AWS Glue documentation on scaling AWS Glue job resources: https://docs.aws.amazon.com/glue/latest/dg/monitor-profile-glue-job-cloudwatch-metrics.html
By checking these documentation resources, the data engineer can get insight into AWS best practices and recommendations for optimizing AWS Glue jobs, which supports the suggested actions to address the slow job performance.
2
Here you can find 5 different Worker types:
https://docs.aws.amazon.com/glue/latest/dg/add-job.html
2
It seems there are actually different worker types in AWS Glue. I'll go with AB as well.
"With AWS Glue, you only pay for the time your ETL job takes to run. There are no resources to manage, no upfront costs, and you are not charged for startup or shutdown time. You are charged an hourly rate based on the number of Data Processing Units (or DPUs) used to run your ETL job. A single Data Processing Unit (DPU) is also referred to as a worker. AWS Glue comes with three worker types to help you select the configuration that meets your job latency and cost requirements. Workers come in Standard, G.1X, G.2X, and G.025X configurations."
https://docs.aws.amazon.com/glue/latest/dg/components-key-concepts.html
1
I would also pick A, B.
But there are no worker types in AWS Glue. You can only raise the DPU.
1
How does partitioning data in S3 help improve the performance of AWS Glue jobs? Partitioning data s3 improves query performance, but the question was what action the DE should take to improve the performance of AWS Glue jobs !