ETExamTower
Q18Data Ingestion and TransformationMultiple answers

A data engineer is building a data pipeline on AWS by using AWS Glue extract, transform, and load (ETL) jobs. The data engineer needs to process data from Amazon RDS and MongoDB, perform transformations, and load the transformed data into Amazon Redshift for analytics. The data updates must occur every hour. Which combination of tasks will meet these requirements with the LEAST operational overhead? (Choose two.)

Select 2 answers.
← → navigate · a answer
Community votes
A
45% (9)
D
40% (8)
C
10% (2)
B
5% (1)
E
0% (0)
Discussion · 12
A, D 7
Selected Answer: AD AWS Glue triggers offer a simple, integrated way to schedule ETL jobs. By setting these triggers to run hourly, the data engineer can make sure the data processing and updates happen as needed without external scheduling tools or custom scripts. This is directly integrated with AWS Glue, which lowers complexity and operational overhead. AWS Glue supports connections to different data sources, including Amazon RDS and MongoDB. Using AWS Glue connections, the data engineer can easily set up and manage connectivity between these data sources and Amazon Redshift. This uses AWS Glue’s built-in data source integration capabilities, helping minimize operational complexity and providing a smooth data flow from the sources to the destination (Amazon Redshift).
A, D 6
Selected Answer: AD A. Set up AWS Glue triggers to run the ETL jobs every hour. Reduced Code Complexity: Glue triggers remove the need to write custom code for ETL job scheduling. This keeps the pipeline simpler and lowers maintenance overhead. Scalability and Integration: Glue triggers integrate smoothly with Glue ETL jobs, ensuring efficient scheduling and execution inside the Glue ecosystem. D. Use AWS Glue connections to set up connectivity between the data sources and Amazon Redshift. Pre-Built Connectors: Glue connections provide pre-built connectors for several data sources like RDS and Redshift. This avoids manual setup and makes data source access inside the ETL jobs simpler. Centralized Management: Glue connections are centrally managed in the Glue service, which streamlines connection management and cuts operational overhead.
A, D 4
Selected Answer: AD A - this is obvious, and D - https://docs.aws.amazon.com/glue/latest/dg/console-connections.html
3
A. Configure AWS Glue triggers to run the ETL jobs every hour. D. Use AWS Glue connections to establish connectivity between the data sources and Amazon Redshift. Explanation: Option A: Configuring AWS Glue triggers lets the ETL jobs be scheduled and run automatically every hour without manual intervention. This lowers operational overhead by automating the data processing pipeline. Option D: Using AWS Glue connections simplifies connectivity between the data sources (Amazon RDS and MongoDB) and Amazon Redshift. Glue connections hide the details of connection configuration, making the data pipeline easier to manage and maintain.
A, D 3
Selected Answer: AD Not a clear question - B would kind of make sense - but AD seems more correct
A, D 3
Selected Answer: AD I was not sure about A - but in AWS console => Glue => Triggers => Add Trigger, I found the Trigger type: "Schedule - Fire the trigger on a timer."
A, B 2
Selected Answer: AB Lambda triggers for Glue jobs still make me dizzy
A, D 2
Selected Answer: AD A. AWS Glue has a built-in way to trigger ETL jobs on scheduled intervals, like every hour. Using Glue triggers reduces the need for extra custom code or services, which lowers operational overhead. D. AWS Glue connections make it easier to establish secure and reliable connections to various data sources (Amazon RDS, MongoDB) and the destination (Amazon Redshift). This reduces the need to manually configure connection settings and makes the ETL pipeline simpler to maintain.
C, D 1
Selected Answer: CD I found this question pretty confusing. At which step would the transformation itself be implemented? I may be wrong, but with Glue triggers we would only run the job, not the transformation logic itself. In this case, I would go with C and D
1
AE D is not valid. It should be Use AWS Glue connections to establish connectivity between the data sources (including Amazon Redshift) and Glue Job
1
An AWS Glue connection is a setting that lets an AWS Glue job access a data source. This lets you connect to databases such as RDS, MongoDB, etc. However, this comment says that the connection is not used to load data directly into Redshift, and that Glue jobs must use the COPY command to load data into Redshift, which is not appropriate. Even so, since Glue jobs can process data and load it directly into Redshift, it is a bit of a stretch to call option D unconditionally wrong.
A, C 1
Selected Answer: AC A. because the question says that the jobs are built in Glue, and they must run every hour. C. because you can run the jobs as Lambda functions every hour. B. discarted, because the question says that "DE" is using Glue, DataBrew is for cleaning data without code, but it seems that "DE" is writing code for transforming the data. D. Discarted, because the connections are not directly related to the question, which says that you should run Glue jobs every hour, and the connections don't seem relevant. E. Discarted, because it says that the data source is RDS and MongoDB, not Redshift, so you cannot use the Redshift Data API to get the data and transform it.