Q93Design Cost-Optimized Architectures
A company is building machine learning (ML) models on AWS as independent microservices. At startup, the microservices retrieve approximately 1 GB of model data from Amazon S3 and load it into memory. Users access the ML models through an asynchronous API and can submit either one request or a batch of requests. The company offers the ML models to hundreds of users. Model usage is irregular: some models go unused for days or weeks, while others receive batches of thousands of requests at once. Which solution will satisfy these requirements?
← → navigate · a answer
Community votes
Discussion · 6
D 2
Given that the problem statement does not specify the processing time, Option D (ECS) is the safer choice because it does not have the same limitations as Lambda(cold start and 15 ,minutes execution). However, if the processing time is short and cold starts are acceptable, Option C (Lambda) could also be a cost-effective solution.
C 2
Since AWS Lambda supports 1GB of memory and can be scaled seamlessly, the ans should be C. Question says a lot of APIs are not used frequently so keeping them in ECS will result in higher costs. Also these are async operations means the response is not time bound hence we can wait for lambda to startup and scale up based on the size of the SQS queue.
B 2
Why should I use SQS in option D? Wouldn't ALB be enough?
1
It's talking about accessing models through an asynchronous API, so decoupling is needed (SQS)
D 1
A - Lambda has a cold start latency. There are ways to optimize that (like the environment variable config which I actually did when I used Lambda once), but loading 1GB into memory is still too much.
B - Using ALB makes handling sudden bursts of traffic even more challenging.
C - Just like A if I'm correct.
D - Each ML model needs to load 1 GB of model data at startup ==> Containers in ECS are well-suited for this because they provide persistent execution environments that avoid the overhead of reloading the model data frequently.
D 1
SQS + ECS Auto Scaling is the classic AWS pattern for asynchronous, bursty workloads.