BeHive is an open-source research engine that converts any topic into structured, scored claims, entity graphs, and synthesized reports using multi-layered web scraping and LLM extraction.

Stars

64

7-day growth

No data

Forks

9

Open issues

0

License

MIT

Last updated

2026-07-29

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It produces machine-readable intelligence (scored claims, knowledge graphs) rather than unstructured text, uses a stealth drone arsenal to bypass anti-bot defenses and paywalls, and integrates natively with AI assistants via MCP, all under an MIT license with self-hosting.

Who it is for

  • AI developers building agentic research tools
  • Data scientists and analysts needing verifiable, structured data
  • Journalists and researchers who require multifaceted source gathering
  • Developers integrating research capabilities into existing workflows

Use cases

  • Enhancing AI assistants (Claude, ChatGPT) with real-time, cited research
  • Competitive and market intelligence gathering from 70+ specialized APIs
  • Academic literature review with automatic claim extraction and quality scoring
  • Building persistent knowledge graphs that accumulate across research sessions

Strengths

  • Dual-model extraction (fast Haiku for bulk, powerful Sonnet for enrichment) with per-claim quality scoring (0.55 threshold)
  • Stealth drone fetch stack (8 layers) that bypasses Cloudflare, DataDome, paywalls, and rate limits
  • 70+ API sources across 37 categories (academic, financial, government, security, etc.) for broad coverage
  • Native MCP support plus REST/SSE streaming for real-time progress and integration with any MCP client

Considerations

  • Requires your own LLM API key (costs $0.30–$2.00 per research mission in tokens)
  • Full feature set relies on PostgreSQL, Neo4j, and optional Qdrant (self-hosted infrastructure)
  • Stealth scraping may violate website terms of service; ethical use is the user's responsibility

README quick start

Quick Start

Description

Open-source deep research engine that builds structured knowledge graphs. MCP-native. avg 0.82+ quality.

Related repositories

Similar projects matched by category, topics, and programming language.

0xwilliamortiz
Featured
0xwilliamortiz GitHub avatar

openclaude-improved

OpenClaude is an open-source CLI coding agent that runs on any platform and supports a wide range of LLM providers, offering the same tools and workflows as Claude Code.

AI & Machine LearningLarge Language Models
577
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
gavamedia
Featured
gavamedia GitHub avatar

deltafin

Deltafin is a research project that runs the 2.8-trillion-parameter Mixture-of-Experts model Kimi K3 on a single Apple Silicon Mac (e.g., M1 Max with 64 GB) at about 16 seconds per token, using exact, reproducible inference with local or streaming expert loading.

AI & Machine LearningLarge Language Models
304