Q14Data Ingestion and Transformation
A company stores daily records of the financial performance of investment portfolios in .csv format in an Amazon S3 bucket. A data engineer uses AWS Glue crawlers to crawl the S3 data. The data engineer must make the S3 data accessible daily in the AWS Glue Data Catalog. Which solution will meet these requirements?
← → navigate · a answer
Community votes
Discussion · 9
11
B. Create an IAM role that includes the AWSGlueServiceRole policy. Attach the role to the crawler. Set the S3 bucket path of the source data as the crawler's data store. Create a daily schedule to run the crawler. Set a database name for the output.
Explanation:
Option B correctly sets up the IAM role with the required permissions by using the AWSGlueServiceRole policy, which is intended for AWS Glue. It sets the S3 bucket path of the source data as the crawler's data store and creates a daily schedule to run the crawler. It also sets a database name for the output, making sure the crawled data is properly cataloged in the AWS Glue Data Catalog.
B 7
Selected Answer: B
A,C are wrong because you don't need full S3 access. D is wrong because you don't need to provision DPU and the destination should be a database, not an s3 bucket. so it's B
B 5
Selected Answer: B
Glue Crawlers are serverless. The point where i saw Assigning DPUs was when I decided it option B
3
S3 access is included in the AWSGlueServiceRole Policy
https://docs.aws.amazon.com/aws-managed-
policy/latest/reference/AWSGlueServiceRole.html
B 3
Selected Answer: B
answer B is incomplete. Even if we include AWSGlueServiceRole policy on IAM role, S3 access is not garantee
1
How does Glue get access to S3 if you don't do B?
1
I meant A
1
It adds only for the glue related buckets, but it doesnt grant permissions for S3 that we need to read in order to fetch data, isnt it ?
B 1
Selected Answer: B
Glue crawlers require the AWSGlueServiceRole policy. Run the crawler daily and specify the Data Catalog database. No need to set up S3 output or DPUs.