ETExamTower
Q14ML Model Development

Case study - An ML engineer is developing a fraud detection model on AWS. The training dataset includes transaction logs, customer profiles, and tables from an on-premises MySQL database. The transaction logs and customer profiles are stored in Amazon S3. The dataset has a class imbalance that affects the learning of the model's algorithm. Additionally, many of the features have interdependencies. The algorithm is not capturing all the desired underlying patterns in the data. The ML engineer needs to use an Amazon SageMaker built-in algorithm to train the model. Which algorithm should the ML engineer use to meet this requirement?

← → navigate · a answer
Community votes
A
50% (9)
B
50% (9)
C
0% (0)
D
0% (0)
Discussion · 20
A 11
Selected Answer: A https://docs.aws.amazon.com/en_kr/sagemaker/latest/dg/lightgbm.html
B 7
Selected Answer: B Answer is B
A 5
Selected Answer: A Here's why LightGBM is the best fit for this fraud detection task: Handling Class Imbalance: LightGBM is especially effective at handling imbalanced datasets, which is a key issue mentioned in the problem statement. It has built-in mechanisms to deal with class imbalance. Feature Interdependencies: LightGBM can capture complex feature interactions through its tree-based structure, addressing the issue of feature interdependencies mentioned in the problem. Capturing Underlying Patterns: As an advanced gradient boosting framework, LightGBM is excellent at capturing complex patterns in data, which the current algorithm is struggling with. Suitable for Fraud Detection: LightGBM is widely used in fraud detection tasks due to its high performance and ability to handle large datasets efficiently. Handling Various Data Types: It can work well with the mix of data types likely present in transaction logs, customer profiles, and database tables.
A 4
Selected Answer: A We have an unbalanced dataset, so this means we have a labelled dataset and are going to use supervised model training. This reduces the options to A and B (K-means and NTM are unsupervised). Both LightGBM and Linear Learner provide hyperparameters to manage unbalanced datasets, respectively "scale-pos_weight" and "positive_example_weight_mult". I would go for LightGBM as this algorithm is more suited to handle complex relationships among features, while Linear Learner learns a linear function, or, for classification problems, a linear threshold function, and maps a vector x to an approximation of the label y.
A 3
Selected Answer: A This is a binary classification problem, so LightGBM should be used. The other algorithms are not for binary classification.
A 3
Selected Answer: A https://docs.aws.amazon.com/en_kr/sagemaker/latest/dg/lightgbm.html
A 3
Selected Answer: A Is supported by Sagemaker
A 3
Selected Answer: A ChatGPT says it's A: LightGBM
A 3
Selected Answer: A 1. Clase desbalanceada: LightGBM (Light Gradient Boosted Machine) es muy efectivo para trabajar con datasets desbalanceados, gracias a su capacidad para ajustar los pesos de clases y usar técnicas como weighted loss o boosting adaptativo. 2. Interdependencia entre variables: LightGBM puede capturar relaciones no lineales e interacciones entre variables gracias a su estructura basada en árboles de decisión. Esto lo hace más apropiado que modelos lineales como Linear Learner, que solo capta relaciones lineales. 3. No captura de patrones complejos: La descripción indica que el algoritmo actual no está capturando patrones complejos subyacentes, lo que sugiere la necesidad de un modelo más robusto como LightGBM, capaz de modelar relaciones complejas y no lineales en los datos
B 2
Selected Answer: B A. Light BGM : It's a suitable model, but not a built-in model for SageMaker. Answer B. Linear learner : suitable model, built-in model for SageMaker. C. K-means clustering : groups similar data points, not suitable for classification problems, and it's an unsupervised learning algorithm so it doesn't fit in this case (fraud detection). D. Neural Topic Model: used for topic modeling and document classification, not suitable for fraud detection
B 2
Selected Answer: B Linear Learner is a built-in algorithm from SageMaker for supervised learning tasks like regression and classification. LightGBM is not a built-in algorithm in Amazon SageMaker. While it is a strong gradient-boosting algorithm, it would need to be implemented as a custom script in SageMaker, which adds operational overhead.
2
LightGBM is better for this use case https://docs.aws.amazon.com/sagemaker/latest/dg/lightgbm.html
B 2
Selected Answer: B Fraud detection is a binary classification problem, and Linear Learner is built for classification tasks. Although LightGBM can handle binary classification tasks, including fraud detection, and it is actually widely used for fraud detection, it is not available as a built-in SageMaker algorithm. Linear Learner is a built-in SageMaker algorithm.
B 1
Selected Answer: B Linear Learner. LightGBM is NOT a built-in algorithm, which the question asks for.
B 1
Selected Answer: B In an ideal situation, for a problem with these traits (fraud detection, class imbalance, feature interdependencies, complex patterns), a tree-based ensemble method like XGBoost (which is a SageMaker built-in algorithm) would be more appropriate. XGBoost can handle non-linear relationships, is robust to class imbalance with proper tuning, and can capture complex patterns in the data. However, given the options provided and the requirement to use a SageMaker built-in algorithm, the Linear learner is the best available choice among these options for this specific fraud detection task.
1
Light BGM is built-in model for SageMaker https://docs.aws.amazon.com/sagemaker/latest/dg/lightgbm.html
1
LightGBM is built-in https://docs.aws.amazon.com/sagemaker/latest/dg/lightgbm.html
B 1
Selected Answer: B Linear learner is a built-in algorithm where LightBM is not
B 1
Selected Answer: B Linear learner is a built-in algorithm, whereas LightGBM can be used through a custom container.
B 1
Selected Answer: B because linear learner is a built-in algorithm while lgbm is not