moondevonyt GitHub avatar

Moon-Dev-AI-Trading-Battles

moondevonyt

Six flagship AI models trade real $100 Hyperliquid perpetuals accounts against each other hourly, with all decisions and trades publicly recorded on-chain to benchmark which model is best at trading.

Stars

11

7-day growth

No data

Forks

4

Open issues

0

License

No data

Last updated

2026-07-27

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It's a live, transparent, real-money benchmark that pits top AI models head-to-head under identical conditions, with every prompt, decision, and trade publicly auditable in an open-source repo.

Who it is for

  • AI researchers and model developers
  • Crypto traders and quant enthusiasts
  • Investors evaluating AI trading capabilities
  • Open-source and transparency advocates

Use cases

  • Comparing decision-making performance of frontier AI models in a real market
  • Evaluating AI risk management and execution without human intervention
  • Providing a public benchmark for AI trading that can be independently verified
  • Understanding how different models behave under identical market conditions over many cycles

Strengths

  • Fully open-source and on-chain: every prompt, decision, and trade is verifiable
  • Real money ($100 each) and real risk, not paper trading
  • Controlled experiment: identical market snapshot, rules, and execution for all models
  • Elimination rule ensures underperformers are removed and don't skew long-term results

Considerations

  • Not financial advice and explicitly not plug-and-play for personal use
  • No stop-losses or take-profits means high risk of total loss for underperforming models
  • Limited to one asset (BTC) and one exchange (Hyperliquid) with fixed 1x leverage

README quick start

🥊 Moon Dev's AI Trading Battles

Live at moondev.com/ai. Six flagship AI models each trade a real $100 Hyperliquid perpetuals account against each other. Every hour they all get the exact same market snapshot and decide: long, short, or flat. No human help, no algorithm behind it, the model's decisions only. Every decision is public forever.

The question this exists to answer

A new frontier model ships what feels like every month now, from a different lab every time, each one benchmarked to death on math, code, and trivia. None of that answers the only question a trader cares about:

Which AI is actually the best at trading?

Not which one predicts best. Which one decides best: sizing, timing, when to sit still, when to cut.

So they all get the identical data, the identical rules, and 1,000 decisions to prove it. Same snapshot down to the tick, same prompt, same market, same hour. Whatever separates them is the model.

Predictions and hot streaks aren't skill. Execution, risk, and time are everything. This is the running, public proof, the benchmark that gets re-run every time a lab ships a new flagship.

🔍 100% transparent, on purpose

A benchmark you can't audit is a marketing page. So:

  • The whole harness is this repo. Every prompt, the exact market snapshot, the cadence, the position sizing, the execution path, all of it is right here, readable. Nothing about how these models are asked or how their trades are placed is hidden.
  • Every trade settles on chain. On moondev.com/ai you can click any model to open its live P&L curve and its actual Hyperliquid address, and watch its positions in real time. You don't have to trust the leaderboard, go read the chain.
  • Every decision is posted the moment it's made, with the model's complete unedited reasoning. Including the bad ones. Especially the bad ones.

The Fighters

FighterLabModel (via OpenRouter)
CLAUDEAnthropicanthropic/claude-opus-5
OPENAIOpenAIopenai/gpt-5.6-sol
GEMINIGooglegoogle/gemini-3.1-pro-preview
GROKxAIx-ai/grok-4.5
KIMIMoonshotmoonshotai/kimi-k3
DEEPSEEKDeepSeekdeepseek/deepseek-v4-pro

The Bench: waiting labs (Meta, Z.ai/GLM, MiniMax, Qwen…), ranked by independent intelligence index, prom

Description

🥊 Six frontier AI models each trade a real $100 Hyperliquid account on identical hourly data. The open benchmark for which AI actually trades best. Live at moondev.com/ai

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482