A local audio processing tool that transcribes speech and analyzes prosody (pitch, energy, pauses) using a biomimetic cochlea approach, enabling AI models to understand both content and emotional tone of voice.

Stars

17

7-day growth

No data

Forks

3

Open issues

0

License

MIT

Last updated

2026-07-30

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It bridges the gap between audio and text-based AI by converting speech into visual and numerical representations that multimodal models can interpret, all while keeping data local.

Who it is for

  • Developers building AI companions
  • Voice interaction researchers
  • Privacy-conscious users wanting local voice processing
  • Chatbot creators integrating emotional nuance

Use cases

  • Empower AI chatbots to respond to emotional tone in voice messages
  • Analyze speech patterns for prosody research
  • Enable privacy-first voice transcription and analysis
  • Augment language models with non-verbal communication cues

Strengths

  • Fully local processing ensures data privacy
  • Biomimetic cochlea method extracts meaningful prosodic features
  • Produces both visual (spectrogram) and numeric summaries
  • Combines transcription with prosody analysis in one pipeline

Considerations

  • Requires ffmpeg for non-WAV audio formats
  • First-time run downloads Whisper model (140MB+)
  • Designed for file input, not real-time streaming

README quick start

安装

Description

给AI伴侣装耳朵——本地语音转写+韵律分析

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482