Q31Data Ingestion and Transformation
A data engineer needs to use AWS Step Functions to design an orchestration workflow. The workflow must parallel process a large collection of data files and apply a specific transformation to each file. Which Step Functions state should the data engineer use to meet these requirements?
← → navigate · a answer
Community votes
Discussion · 8
C 3
Selected Answer: C
C is Correct
To meet the requirement of parallel processing a large collection of data files and applying a specific transformation to each file, the data engineer should use the Map state in AWS Step Functions.
The Map state is specifically built to run a set of tasks in parallel for each element in a collection or array. Each element (in this case, each data file) is processed separately and in parallel, letting the workflow take advantage of parallel processing.
C 3
Selected Answer: C
The Map state lets you define one execution path for processing a collection of data items in parallel.
This fits perfectly with the data engineer's need to parallel process a large collection of data files
C 2
Selected Answer: C
The Map state is specifically built for processing a collection of items (like data files) in parallel. It lets you apply a transformation or a set of steps to each item in the input array independently.
The Map state automatically iterates through each item in the array and carries out the defined steps. This makes it ideal for cases where you need to process a large number of files in a similar way, as in your requirement.
1
C is correct.
Map state is built exactly for the requirement described. It lets you iterate over a collection of items, processing each item one by one. The Map state can automatically handle the iteration and run the specified transformation on each item in parallel, making it the ideal choice for parallel processing of a large collection of data files.
1
With Step Functions, you can orchestrate large-scale parallel workloads to perform tasks, such as on-demand processing of semi-structured data. These parallel workloads let you process large-scale data sources stored in Amazon S3 at the same time. For example, you might process a single JSON or CSV file that contains large amounts of data. Or you might process a large set of Amazon S3 objects.
To set up a large-scale parallel workload in your workflows, add a Map state in Distributed mode.
C 1
Selected Answer: C
C, Map state is the correct choice
C 1
Selected Answer: C
to run in parallel
C 1
Selected Answer: C
Clearly this is the mapping state