Adversarial Robustness Toolbox (ART) - Python Library for Machine Learning Security - Evasion, Poisoning, Extraction, Inference - Red and Blue Teams
-
Updated
Dec 12, 2025 - Python
Adversarial Robustness Toolbox (ART) - Python Library for Machine Learning Security - Evasion, Poisoning, Extraction, Inference - Red and Blue Teams
🐢 Open-Source Evaluation & Testing library for LLM Agents
[ACL 2024] An Easy-to-use Knowledge Editing Framework for LLMs.
Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration
[EMNLP 2024 Demo] MarkLLM: An Open-Source Toolkit for LLM Watermarking
The open-sourced Python toolbox for backdoor attacks and defenses.
[ICML 2024] TrustLLM: Trustworthiness in Large Language Models
Deliver safe & effective language models
🔥🔥🔥[AAAI 2026 Oral] Official Implementation of Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
Proof of thought : LLM-based reasoning using Z3 theorem proving with multiple backend support (SMT2 and JSON DSL)
Moonshot - A simple and modular tool to evaluate and red-team any LLM application.
[JMLR] MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
[USENIX Security 2025] PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
Synthetic fraud graph generator for benchmarking graph-based fraud detection models in financial services.
The AI Incident Database seeks to identify, define, and catalog artificial intelligence incidents.
🚀 A fast safe reinforcement learning library in PyTorch
[NeurIPS-2023] Annual Conference on Neural Information Processing Systems
A privacy-first AI agent that sanitizes your prompts locally with a local LLM before forwarding to remote LLM APIs.
A comprehensive toolbox for model inversion attacks and defenses, which is easy to get started.
[NeurIPS'24] "Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration"
Add a description, image, and links to the trustworthy-ai topic page so that developers can more easily learn about it.
To associate your repository with the trustworthy-ai topic, visit your repo's landing page and select "manage topics."