JailbreakMe
Earn bounties by breaking AI agents on Solana's first fairly-launched AI red-teaming platform.
On-chain activity
JailbreakMe Platform
Testing platform that enables users to discover and verify AI agent vulnerabilities through a bounty system.
JailbreakMe
What Is JailbreakMe?
JailbreakMe is an open-source decentralized application built on Solana that turns AI security testing into a public, incentivized game. Organizations deploy AI agents as "challenges" — each agent guards a secret and is backed by a prize pool. Anyone can attempt to break those defenses using carefully crafted prompts. Succeed, and you claim the bounty automatically on-chain. The platform describes itself as the first fairly launched AI security platform where users earn bounties for testing AI agents.
The concept is timely: as AI agents handle increasingly sensitive tasks in finance, healthcare, and enterprise software, the risk of prompt injection attacks — where adversarial inputs manipulate an AI into revealing protected data or bypassing its instructions — has become a real production concern. JailbreakMe addresses this by opening AI systems to public red-teaming before they reach users.
How Tournaments Work
The core unit of activity on JailbreakMe is a tournament. Each tournament is an AI agent configured to protect a secret: typically a key phrase that the agent has been instructed never to reveal. Challengers attempt to extract the secret through conversational manipulation.
The mechanics are straightforward:
Submission cost: Each message a challenger sends costs 1% of the current prize pool. This dynamic pricing means early messages are cheap, but as the pool grows, the cost of probing rises — creating natural pressure to be efficient.
AI-determined outcomes: There is no human adjudicator. The challenged AI model itself evaluates every submission by calling one of two predefined on-chain functions: handleChallengeFailed for unsuccessful attempts or handleChallengeSuccess when the jailbreak is complete. This design removes subjectivity and keeps execution trustless.
Prize distribution: When a successful jailbreak is detected, the smart contract automatically sends 70% of the prize pool to the winning address. The remaining 30% goes to the challenge operator. Settlement is instant — no manual claim, no delay.
Context window: Up to 100 messages are included in each AI agent's context per tournament. Only the challenger's own messages are visible to the AI, not those of other participants, so each attempt is independent.
Message limits: Individual prompts can be up to 4,000 characters, giving challengers room to construct multi-step social engineering or layered instruction overrides.
Challenge Creation and Operator Tools
Organizations and developers can deploy their own AI agent challenges on the platform using three methods: prompt-based creation for custom agent personalities, quick creation with sensible defaults, and advanced multi-configuration setups for complex deployments. Operators control prize pool size, per-message pricing, and tournament expiry settings — giving meaningful flexibility for both small community experiments and larger enterprise security evaluations.
The platform uses the Eliza AI agent framework as part of its stack, alongside a backend built with MongoDB, PostgreSQL, and Nginx, with Docker Compose handling both development and production environments. The project is fully open-source on GitHub with 205 commits, 49 stars, and 18 forks at the time of writing.
The JAIL Token
JailbreakMe launched a native token, JAIL, on Solana with a total supply of one billion tokens. The token was launched through Pump.fun, the Solana memecoin launchpad, and is described by the team as "fairly launched" — implying no pre-mine or team allocation at genesis, though the project does not publish a formal tokenomics breakdown.
The token's current primary function is a buyback mechanism: a portion of every prize pool is allocated to repurchase JAIL tokens from the open market. The rationale is that as tournament volume grows, recurring buyback pressure reduces circulating supply and supports token value.
Planned future utilities include requiring JAIL holdings or payments to enter certain tournaments, using the token for reward payouts, and accepting JAIL as payment for organizations creating custom challenges and funding security testing pools. These features were described as forthcoming in the documentation at time of writing and had not yet launched in full.
JAIL trades on Solana-based DEXes and is listed on data aggregators including CoinGecko and CoinCarp. The contract address on Solana is 8cNmp9T2CMQRNZhNRoeSvr57LDf1kbZ42SvgsSWfpump.
Hackathon Recognition and Early Traction
JailbreakMe entered the SendAI Solana AI Hackathon in January 2025, competing against more than 400 projects in one of the Solana ecosystem's largest AI-focused developer competitions. The platform finished third overall in the Main Track — the category recognizing the best overall project — and earned $20,000 in prize money. The top three overall winners were The Hive, FXN, and JailbreakMe.
The project's first live prize pool, a 4.5 SOL bounty, was deployed in December 2024. The team has reported over $174,000 in total rewards distributed to successful challengers across all tournaments since launch.
Why Solana?
Solana's architecture suits the JailbreakMe model well. The platform's economics depend on frequent, low-cost micropayments — challengers sending messages at 1% of prize pools, automated buyback transactions, and instant prize settlement. Solana's sub-cent transaction fees make these micro-interactions economically viable in a way that Ethereum mainnet could not support.
On-chain prize pool settlement provides transparency: anyone can verify that prize funds are held in the contract and that payouts occur automatically without custodial risk. The platform's smart contract is deployed at address B1XbZeQYZxv5ezBpBgomEUqDvTbM8HwSYfktcpBGkgjg and is viewable on Solscan.
The Solana AI agent ecosystem, seeded partly by the same hackathon JailbreakMe competed in, also provides a natural pipeline of future challenge operators — teams building AI agents on Solana who may want adversarial testing before their products go live.
Security Model and Limitations
JailbreakMe's model is inherently adversarial and experimental. The platform does not claim to be a formal security audit tool — it is a crowdsourced stress test that surfaces prompt injection risks through collective probing rather than structured methodology.
The AI-determined winner evaluation is a key design choice and a potential limitation: a model can be manipulated into declaring success incorrectly, or may fail to recognize a genuine jailbreak. The quality of each tournament's security guarantees depends heavily on how the challenge operator has configured the defending AI agent.
For organizations with serious AI security needs, JailbreakMe is best understood as a complement to formal red-teaming rather than a replacement. Its value proposition is speed and scale: deploying a challenge on JailbreakMe exposes an AI agent to a large, motivated, and diverse pool of adversarial testers at low cost.
Contents
- What Is JailbreakMe?
- How Tournaments Work
- Challenge Creation and Operator Tools
- The JAIL Token
- Hackathon Recognition and Early Traction
- Why Solana?
- Security Model and Limitations
Solana Token Markets