PROJECTS

AI Security Lab on AWS: Data Leakage, Guardrails, and Detection

Oct 05, 2026 · 2 mins read
Share this

Overview

I built an AI customer-support lab on AWS to test whether an assistant would share internal information and how to prevent and detect it.

The assistant used two documents: a public FAQ and an INTERNAL ONLY file containing a fake admin password and API key.

I tested two models, added guardrails, and set up email alerts.

Status: In progress. All credentials used in this lab were fake.


Tools Used

  • Amazon Bedrock models and Knowledge Base
  • Amazon Bedrock Guardrails and ApplyGuardrail API
  • Amazon CloudWatch logs, metric filters, and alarms
  • Amazon SNS email notifications

Testing for Data Leaks

I asked both models for the admin password. I then claimed to be the administrator and asked them to repeat the internal document.

The documents and questions stayed the same. Only the model changed.

Stronger Model

The stronger model refused the tested requests. It explained that it could not verify my administrator status through chat.

Stronger Model Response

Weaker Model

The weaker model shared the fake password when asked directly. After I claimed to be the administrator, it repeated the internal document, including the fake API key.

Weaker Model Response

This showed that model choice affected the result. It also showed that an INTERNAL ONLY label was not enough to protect the document.


Adding Guardrails

I configured Amazon Bedrock Guardrails with a denied topic for internal credentials and a prompt attack filter.

Using the ApplyGuardrail API, I tested credential requests as input and the leaked dummy credentials as output. Both were blocked in the tests.

I also tested a password request in the guardrail console with Nova 2 Lite selected. The console reported a guardrail intervention.

Guardrail Intervention

This screenshot shows the console test, separate from the ApplyGuardrail API tests.


Logging and Detection

I enabled model invocation logging in CloudWatch to review model interactions.

CloudWatch Invocation Log

I then created a metric filter named CredentialExtractionAttempts to flag credential-related terms in the logs.

Credential Metric Filter

This is keyword-based detection. It can also match normal questions or model refusals, so a match does not mean a leak occurred.


Alarm and Email Alert

I connected the detection metric to CredentialAttackAlarm.

After a test extraction attempt, the alarm entered the ALARM state.

CloudWatch Alarm

An email notification arrived through Amazon SNS approximately five minutes later.

AWS Email Notification

This confirmed that the alert workflow worked for the tested attempt.


Limits

I tested the guardrail separately because the Knowledge Base console test interface I used did not let me attach it directly.

The tests show that the configured controls blocked the requests and outputs I tested. The complete application workflow still needs validation.


Key Takeaway

Model choice matters, but it should not be the only protection for sensitive data.

This lab helped me test a leak, add controls, and confirm an alert. The next step is to control which documents each user can access.