Q6Security, Compliance, and Governance for AI Solutions
An AI practitioner trained a custom model on Amazon Bedrock by using a training dataset that contains confidential data. The AI practitioner wants to ensure that the custom model does not generate inference responses based on confidential data. How should the AI practitioner prevent responses based on confidential data?
← → navigate · a answer
Community votes
Discussion · 13
A 6
To ensure that the custom model does not generate inference responses based on confidential data, the best approach is to:
Delete the custom model: If the confidential data was used in training, there's a possibility that the model has memorized this data and might generate it in responses. Removing the model is the first step.
Remove the confidential data from the training dataset: This ensures that confidential information is not included in the model's learning process, mitigating the risk of leakage.
Retrain the custom model: After removing the confidential data, retraining the model with a cleaned dataset ensures that the model does not inadvertently include any sensitive information in its responses.
A 5
A: Delete the custom model. Remove the confidential data from the training dataset. Retrain the custom model.
Explanation:
If the training dataset contains confidential data, the model may inadvertently learn and generate responses based on that data. The only way to ensure that the model does not generate responses based on the confidential data is to:
Remove the confidential data from the training dataset.
Retrain the custom model using the updated dataset.
This process ensures that the model is not influenced by the sensitive information.
A 3
Explanation:
Once a model is trained, the data used for training is embedded in its parameters. If confidential data is included in the training dataset, it can influence the responses the model generates.
Simply masking or encrypting inference responses will not ensure the model doesn’t generate responses derived from the confidential data; the issue originates in the training process itself.
A 3
Delete the custom model, remove the confidential data, and retrain the model is the best approach because it ensures that the model will not retain or generate responses based on any confidential information
B 2
The company should mask the confidential information
A 2
Once a model is trained on data, its outputs may inherently reflect patterns or details derived from the training dataset, including confidential data.
To ensure the custom model does not generate inference responses based on confidential data, the only reliable solution is to:
- Remove the confidential data from the training dataset.
- Retrain the model with the updated dataset.
This approach ensures the model is not influenced by sensitive information during inference.
Option B is incorrect.
Dynamic data masking hides sensitive information in database query results or outputs but does not prevent the model from generating responses influenced by the confidential data. The model would still "know" the sensitive patterns.
A 2
The correct answer is A. Once a model is trained on confidential data, it must be retrained without it.
B 2
This is the most efficient method, effectively maintaining data privacy and security.
A 2
If the model was trained with confidential data, there's a risk it might have memorized that information and could generate it in responses. Deleting the model is the first step to prevent this.
2
A
A. Delete the custom model. Remove confidential data from dataset. Retrain the model.
This is the correct answer because:
Once a model learns from confidential data, that information becomes embedded in its parameters
The only way to truly prevent it from using that knowledge is to retrain from scratch without the confidential data
Deleting and retraining ensures the model has no access to the sensitive information
B 1
Dynamic data masking in Amazon Bedrock can be implemented to protect confidential data within inference responses, particularly when using features like Knowledge Bases and Guardrails for RAG (Retrieval Augmented Generation) applications.
A 1
A is the correct answer
A 1
data memorize is possible.