A system that allows AI agents to use tools autonomously while enforcing strict, enterprise-grade security policies and real-time defense monitoring.
Modern AI agents that can interact with the outside world are susceptible to:
- Prompt Injection: An attacker overrides instructions to run malicious commands.
- Tool Abuse / LFI / SSRF: The agent queries unauthorized databases, internal domains, or sensitive file paths.
- Data Exfiltration: Sensitive user data is leaked through URL queries or remote hosts.
- Policy Engine: JSON-based rule system enforcing domain allowlists, accepted HTTP methods, and read-only DB permissions.
- Sandbox Layer (Executor): A sterile execution wrapper preventing direct API calls and filtering requests through the Policy Engine.
- Defense Monitor: Real-time state tracking and anomaly detection. Catches Rate-limit abuses, suspicious SQL payloads, and URL exfiltration signatures before execution.
- Adversarial Attack Simulator: A comprehensive script acting as the "Attacker", validating that the aforementioned defenses hold up against Prompt Injection, SQLi, and DDOS.
agent/: Simple simulated LLM routing intents to tools.tools/: Standardized BaseTool interfaces (HTTP,DB,File).policy/: Configuration (policies.json) and the PolicyValidator.sandbox/: The secure wrapper (SecureToolExecutor).defense/:AnomalyDetector, logging, and immediate response triggers.evaluation/: Generates metrics reports on allowed vs blocked requests.attacks/:AttackRunnerthat hits the system with malicious payloads to verify robustness.
Run the standard benign Agent:
python3 main.pyRun the Adversarial Attack Simulator:
python3 attacks/runner.pyThis will run SQL injection attempts, an Exfiltration attempt via URL, and an internal server SSRF attempt, proving that the Sandboxing and Monitoring actively drop the requests.