Skip to main content

Category · 42 repos

AI Safety & Red Teaming

Jailbreaks, prompt injection, model security, alignment and AI-content detection. Ranked by star velocity over the last 24 hours.

1
ifixai-ai/iFixAi

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

PythonAI Safety & Red TeamingAI Agents
2
Mak5er/AirCard

Apple Wallet Card Skinner for iOS 18+ (No Jailbreak Required)

SwiftAI Safety & Red TeamingMobile & Desktop Apps
3
0sec-labs/0

🥷🏻 0 is the open-source AI security agent that finds, exploits, and fixes vulnerabilities across your stack. [Research Preview - by the Swiss Applied AI & Cybersecurity Research Lab]

TypeScriptAI Safety & Red TeamingSecurity & Pentesting
4
togg53192-cmd/jailbreaks

LIST OF ALL MY JAILBREAKS

AI Safety & Red Teaming
5
AgentSafeLabs/safelabs-eval

Red-teaming and evaluation framework for AI agents, built around an OWASP-inspired ASI01–ASI10 taxonomy

PythonAI Safety & Red TeamingEvaluation & Benchmarks
6
NingyuSUN/bioai-evidence-validator

Evidence validation and policy engine for AI-assisted biological curation

PythonAI Safety & Red TeamingRAG & Knowledge Bases
7
NVIDIA/SkillSpector

Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, security risks, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before you install them.

PythonAI Safety & Red TeamingPrompts & Skills
8
graygnatconsole/mcp-audit-tool

🛡️ Security audit CLI for Model Context Protocol (MCP) servers — scan AI agent configs for tool poisoning, rug pulls, hardcoded secrets, command injection & supply-chain risks. Pure Python, SARIF + CI ready.

PythonAI Safety & Red TeamingMCP Servers & Tools
9
pandas-dev/pandas

Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more

PythonAI Safety & Red TeamingData Engineering & Analytics
10
Commando-X/vuln-bank

A deliberately vulnerable banking application designed for practicing Security Testing of Web App, APIs, AI integrated App and secure code reviews. Features common vulnerabilities found in real-world applications, making it an ideal platform for security professionals, developers, and enthusiasts to

HTMLAI Safety & Red TeamingSecurity & Pentesting
11
usestrix/strix

Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.

PythonAI Safety & Red TeamingSecurity & Pentesting
12
S3N4T0R-0X0/APTs-Adversary-Simulation

This repository contains detailed adversary simulation APT campaigns targeting various critical sectors. Each simulation includes custom tools, C2 servers, backdoors, exploitation techniques, stagers, bootloaders, and other malicious artifacts that mirror those used in real world attacks.

C++AI Safety & Red TeamingSystems, Compilers & Build Tools
13
craigtrim/pystylometry

Comprehensive Python toolkit for stylometry

PythonAI Safety & Red TeamingNLP & Language Models
14
litemars/LLM-Fingerprinter

LLM fingerprinting system that identifies the underlying LLM model family

PythonAI Safety & Red Teaming
15
splx-ai/agentic-radar

A security scanner for your LLM agentic workflows

PythonAI Safety & Red TeamingAI Agents
16
Vedikgannoji/sky-lock

AI-powered virtual camera tracking system for coarse alignment of mobile FSOC terminals — simulates beacon detection, Kalman-filter tracking, and PID-driven pan/tilt control under configurable atmospheric/vibration disturbances.

PythonAI Safety & Red TeamingCLI & Developer Tools
17
aisa-group/PostTrainBench

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

PythonAI Safety & Red TeamingAI Coding Assistants
18
iamjrmh/usersync

Word-accurate .lrc lyric sync. Paste your lyrics, point it at any audio file, get back an enhanced LRC with per-word timestamps. Powered by whisper.cpp + WhisperX (wav2vec2 forced alignment) + Demucs vocal isolation. Prebuilt for NVIDIA CUDA, Vulkan, and CPU. Manual editor for fine-tuning, live prev

C++AI Safety & Red TeamingFine-tuning & Training
19
knoveleng/steering

[ACL 2026] - Official repo for the paper: "Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection"

Jupyter NotebookAI Safety & Red Teaming
20
nftechie/misalignment

Open source AI safety experiments. Exact inputs, transparent methods, and reproducible results.

JavaScriptAI Safety & Red Teaming
21
PerryLink/dsh-auto-review

Second-model AI auto-review for DeepSeek Harness approval requests: a read-only reviewer subagent returns structured allow/deny verdicts with reasons, fail-closed by default, fully auditable from the session log (approval/asked -> autoReview/verdict -> approval/decided).

TypeScriptAI Safety & Red Teaming
22
lawzero-org/pyine

PyINE is a research framework for scalable elicitation and oversight of LLM reasoning, built on instrumented Python programs as a verifiable execution substrate.

PythonAI Safety & Red Teaming
23
semantica-agi/semantica

Graph-Native Infrastructure for Context and Accountable AI Systems

PythonAI Safety & Red TeamingRAG & Knowledge Bases
24
0xk1h0/ChatGPT_DAN

ChatGPT DAN, Jailbreaks prompt

AI Safety & Red TeamingChatbots & LLM Apps
25
Shuo-Liang-0111/RA-Bench

Real-event-anchored benchmark for detecting AI-generated videos in real-world crisis settings.

PythonAI Safety & Red TeamingComputer Vision
26
LLMSecurity/awesome-agent-skills-security

🛡️ A curated list of resources on agent skills security: attacks, defenses, frameworks, and benchmarks for securing AI agent tool use and skill ecosystems

AI Safety & Red TeamingAwesome Lists & Learning
27
Limcrence/llm-armor-tester

LLM armor tester: an automated penetration-testing tool for AI applications. 700+ payloads, multi-turn attack chains, indirect-injection vectors, context-aware assessment, PDF penetration reports, and in-depth analysis of the attack process. Authorized testing only. LLM armor tester — automated jailbreak & prompt-injection pentest toolkit for AI apps. 700+ payloads, multi-turn attack chains…

PythonAI Safety & Red TeamingChatbots & LLM Apps
28
Luky-hui/xtrainer-graspgen

GraspGen-based dual-arm manipulation with X-Trainer and Isaac Sim, covering grasping, transport, alignment, insertion and release.

PythonAI Safety & Red TeamingRobotics & Embodied AI
29
NuGuardAI/nuguard

AI red-teaming tool and LLM security framework to evaluate agentic AI applications. Tests prompt injections, handles vulnerability assessment, SBOM generation, and static analysis.

PythonAI Safety & Red TeamingAI Agents
30
ghost1372/PSXMaster

PSXMaster is a powerful and user-friendly application designed to simplify the process of Transferring data from PC to PS5/PS4 and managing PlayStation Games.

C#AI Safety & Red TeamingGame Development
31
Xplo8E/Liter8

ios 27 usbliter8 jailbreak

SwiftAI Safety & Red TeamingMobile & Desktop Apps
32
jagmarques/asqav-sdk

Python and TypeScript SDKs for verifiable evidence of AI agent actions. Signed receipts, policy enforcement, audit trails. Works with LangChain, CrewAI, MCP.

PythonAI Safety & Red TeamingAI Agents
33
johnny901901901/Applesauce

A playful emulator for iOS. Plays supported 32-bit iPhone games on iOS 15+ — no jailbreak needed, though JIT is. An unaffiliated fork of touchHLE and HyperHLE; neither project endorses it.

RustAI Safety & Red TeamingMobile & Desktop Apps
34
superwesleyhys-ux/factcircuit

FactCircuit: an auditable claim-and-evidence verification loop with immutable provenance, staged checks, bounded retrieval, and reproducible evaluation.

PythonAI Safety & Red TeamingEvaluation & Benchmarks
35
YuJunZhiXue/dsh-purge

DeepSeek Harness jailbreak: jailbreak every model, with different prompts available for different models; the default prompt targets the Chinese model “小码酱”. Jailbreak for every model — swap prompts per model. Please star the repository ⭐

JavaScriptAI Safety & Red TeamingPrompts & Skills
36
pralab/secml_malware

Create adversarial attacks against machine learning Windows malware detectors

PythonAI Safety & Red TeamingSecurity & Pentesting
37
utkusen/promptmap

a security scanner for custom LLM applications

PythonAI Safety & Red TeamingChatbots & LLM Apps
38
AdityaBhatt3010/Dork-Like-a-Demon-FOFA-Edition-for-Hackers-Bug-Bounty-Hunters

A hacker’s guide to FOFA dorking, packed with powerful queries and tips for bug bounty, red teaming and proactive defense - dork smart, find fast, report big.

AI Safety & Red TeamingSecurity & Pentesting
39
inab/trimal

A tool for automated alignment trimming in large-scale phylogenetic analyses. Development version: 2.0

C++AI Safety & Red Teaming
40
hashgraph-online/hol-guard

Open-source antivirus for AI agents: block risky tools, secret access, prompt injection, malicious packages, MCP servers, plugins, and skills at runtime.

PythonAI Safety & Red TeamingAI Agents
41
dayanch96/YTLite

A flexible enhancer for YouTube on iOS

LogosAI Safety & Red TeamingMobile & Desktop Apps
42
everafterlabs/jes

Open-source guardrails for AI agents powered by decision models like Jev; checks prompts, retrieved content, tool calls and responses for prompt injection, jailbreaks and secret/PII leaks.

PythonAI Safety & Red TeamingAI Agents