Q15Data Operations and Support
A company loads transaction data for each day into Amazon Redshift tables at the end of each day. The company wants to have the ability to track which tables have been loaded and which tables still need to be loaded. A data engineer wants to store the load statuses of Redshift tables in an Amazon DynamoDB table. The data engineer creates an AWS Lambda function to publish the details of the load statuses to DynamoDB. How should the data engineer invoke the Lambda function to write load statuses to the DynamoDB table?
← → navigate · a answer
Community votes
Discussion · 11
B 15
Selected Answer: B
https://docs.aws.amazon.com/redshift/latest/mgmt/data-api-monitoring-events.html
10
The most suitable way for the data engineer to invoke the Lambda function to write load statuses to the DynamoDB table is:
B. Use the Amazon Redshift Data API to publish an event to Amazon EventBridge. Set up an EventBridge rule to invoke the Lambda function.
Explanation:
Option B uses the Amazon Redshift Data API to publish events to Amazon EventBridge, which is a serverless event bus service for handling events across AWS services. By setting up an EventBridge rule to invoke the Lambda function in response to events published by the Redshift Data API, the data engineer can make sure the Lambda function is triggered whenever there is a new transaction data load in Amazon Redshift. This approach gives a simple and scalable solution for tracking table load statuses without depending on extra Lambda functions or services.
B 3
Selected Answer: B
Here's why Option B is the best choice:
Amazon Redshift Data API can be used by applications and scripts to interact with Redshift (for example, run SQL queries, check load status).
Amazon EventBridge can accept custom events or service-generated events and route them to targets like Lambda.
This approach separates the data load process from status logging, using EventBridge as a clean integration point.
An EventBridge rule can filter and trigger the Lambda function when Redshift (or your data pipeline) indicates a successful data load.
2
Im not 100% sure about B or C, this is a tricky question. The reason is that either SQS or EventBridge has no direct native connection to Redshift Data API. There is no way to publish events by itself. So, this means either SQS / EventBridge eventually need a "proxy" (for example lambda function) in order to publish events or process events to this 2 sources. In both services we need something to publish those events from Redshift. so Yes, we need a lamda function between Redshift Data API and (SQS|EB). so either B,C doesnt seem to be 100% right. I think this question its a good candidate to be "Choose two options" but none has 100% right. Both are valid considering that there is an adapter function between 2 solutions.
2
Why not use SQS to keep API changes in the Queue ?
D 2
Selected Answer: D
The statement in B is not accurate.
You don't 'use Amazon Redshift Data API to publish' event to EventBridge. Redshift Data API has no function to write to EventBridge. Instead, the statement should be "Use EventBridge to monitor Data API events..." Maybe this is a typo.
But if I assume there are no typos in any of the statements, then I would choose D. Even though it is not a perfect solution, the CloudTrail events contain more info than the Redshift Data API events.
1
This job doesn’t need a real-time check
1
It seems that Redshift Data API could publish events directly in EventBridge (see first comment). For monitoring the Redshift Data API, you could use either EventBridge (near-real-time) or CloudTrail (stored in S3): https://docs.aws.amazon.com/redshift/latest/mgmt/data-api-monitoring.html
But both services are related to "Data API" and not the Redshift database itself. So it is really tricky.
1
So, you could use the Redshift table STV_LOAD_STATE,
https://docs.aws.amazon.com/redshift/latest/dg/r_STV_LOAD_STATE.html
and run a "select" query on that table to get the status of tables (filtering by timestamp) and add the result to EventBridge, applying a rule on those events to invoke the lambda function. I guess that B is the most appropiate answer.
B 1
Selected Answer: B
the data engineer should use Amazon EventBridge (formerly CloudWatch Events) to trigger the Lambda function based on a schedule or on events that match the completion of the data load process in Amazon Redshift.
B 1
Selected Answer: B
Use the Amazon Redshift Data API to publish an event to Amazon EventBridge.
EventBridge can automatically trigger the Lambda function after each Redshift load event, letting the Lambda write the load status to DynamoDB without manual scheduling or polling.