moondevonyt GitHub avatar

Moon-Dev-AI-Trading-Battles

moondevonyt

六大旗舰AI模型各自使用100美元真实资金在Hyperliquid永续合约上每小时对战,所有决策和交易均公开上链,以衡量哪个模型交易能力最强。

Stars

11

7 天增长

暂无数据

Fork 数

4

开放 Issue

0

开源协议

暂无数据

最近更新

2026-07-27

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

这是一个公开透明的真实资金基准测试,在完全相同的条件下让顶级AI模型直接对决,所有提示、决策和交易均开源可查。

适合谁使用

  • AI研究人员与模型开发者
  • 加密货币交易者与量化爱好者
  • 评估AI交易能力的投资者
  • 开源与透明度倡导者

典型使用场景

  • 比较前沿AI模型在真实市场中的决策表现
  • 评估AI在无人干预下的风险管理和执行能力
  • 提供可独立验证的公开AI交易基准
  • 了解不同模型在相同市场条件下长期的行为差异

项目优势

  • 完全开源且上链:所有提示、决策和交易均可验证
  • 使用真实资金(每模型100美元)和真实风险,非模拟交易
  • 受控实验:所有模型获得完全相同的市场快照、规则和执行
  • 淘汰机制确保表现差的模型被移除,不会扭曲长期结果

使用前须知

  • 非投资建议,明确不适用于个人直接复制使用
  • 无止损或止盈,表现不佳的模型可能面临完全亏损
  • 仅限单一资产(BTC)和单一交易所(Hyperliquid),且杠杆固定为1倍

README 快速开始

🥊 Moon Dev's AI Trading Battles

Live at moondev.com/ai. Six flagship AI models each trade a real $100 Hyperliquid perpetuals account against each other. Every hour they all get the exact same market snapshot and decide: long, short, or flat. No human help, no algorithm behind it, the model's decisions only. Every decision is public forever.

The question this exists to answer

A new frontier model ships what feels like every month now, from a different lab every time, each one benchmarked to death on math, code, and trivia. None of that answers the only question a trader cares about:

Which AI is actually the best at trading?

Not which one predicts best. Which one decides best: sizing, timing, when to sit still, when to cut.

So they all get the identical data, the identical rules, and 1,000 decisions to prove it. Same snapshot down to the tick, same prompt, same market, same hour. Whatever separates them is the model.

Predictions and hot streaks aren't skill. Execution, risk, and time are everything. This is the running, public proof, the benchmark that gets re-run every time a lab ships a new flagship.

🔍 100% transparent, on purpose

A benchmark you can't audit is a marketing page. So:

  • The whole harness is this repo. Every prompt, the exact market snapshot, the cadence, the position sizing, the execution path, all of it is right here, readable. Nothing about how these models are asked or how their trades are placed is hidden.
  • Every trade settles on chain. On moondev.com/ai you can click any model to open its live P&L curve and its actual Hyperliquid address, and watch its positions in real time. You don't have to trust the leaderboard, go read the chain.
  • Every decision is posted the moment it's made, with the model's complete unedited reasoning. Including the bad ones. Especially the bad ones.

The Fighters

FighterLabModel (via OpenRouter)
CLAUDEAnthropicanthropic/claude-opus-5
OPENAIOpenAIopenai/gpt-5.6-sol
GEMINIGooglegoogle/gemini-3.1-pro-preview
GROKxAIx-ai/grok-4.5
KIMIMoonshotmoonshotai/kimi-k3
DEEPSEEKDeepSeekdeepseek/deepseek-v4-pro

The Bench: waiting labs (Meta, Z.ai/GLM, MiniMax, Qwen…), ranked by independent intelligence index, prom

项目描述

🥊 Six frontier AI models each trade a real $100 Hyperliquid account on identical hourly data. The open benchmark for which AI actually trades best. Live at moondev.com/ai

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

lopopolo
精选
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI 与机器学习AI 智能体
2,390
slvDev
精选
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI 与机器学习大语言模型
1,960
littledivy
精选
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI 与机器学习
1,482