VisTouch 是一个公开的机器人滑动接触多模态数据集基准,包含 8 类材料共 2000 个时间对齐的音频/触觉/视频三元组,并提供材质识别、触觉超分辨率、跨模态检索与生成等基线。

Stars

34

7 天增长

暂无数据

Fork 数

3

开放 Issue

0

开源协议

NOASSERTION

最近更新

2026-07-31

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

它提供了大规模同步视觉-触觉-音频的真实传感器数据,并设计了跨压力泛化划分与可复现的轻量基线,是多模态触觉感知研究的重要公开资源。

适合谁使用

  • 机器人触觉与灵巧操作研究人员
  • 多模态机器学习研究者
  • 跨模态语义通信领域人员
  • 材质识别与触觉信号处理开发者

典型使用场景

  • 基于音频、触觉和视频的材质识别模型训练与基准测试
  • 低采样率或噪声触觉信号的超分辨率重建
  • 触觉与声音之间的跨模态检索
  • 仅从声音生成触觉力曲线

项目优势

  • 2000 个真实采集的时间对齐音频/触觉/视频三元组,覆盖 8 类材料
  • 基线效果显著优于随机或朴素方法:材质识别 76.0%,触觉超分辨率相关系数 0.996
  • 按接触力划分训练/测试集,用于评估跨压力泛化能力
  • 代码与数据许可宽松(MIT/CC-BY-4.0),基线可在 CPU 上快速复现

使用前须知

  • 当前公开版本仅包含 47 类材料中的 8 类
  • 数据需从外部网盘下载,仓库内 dataset 目录为空
  • 基线刻意保持轻量,不代表当前最优结果

README 快速开始

A large-scale synchronized vision–touch–audio dataset of robotic sliding contact


🎬 What you can do with VisTouch

  🏷️ Material Recognition
  fused audio+tactile+video · 76.0% test acc (chance 12.5%)


  
  🧵 Tactile Super-Resolution
  clean 100Hz force curve from noisy low-rate input · −85% MSE, 0.996 corr




  
  🔍 Cross-Modal Retrieval
  touch query → matching sound clip · 2× over chance Recall@1/@5


  
  🔄 Cross-Modal Generation
  haptic recovery: tactile force curve generated from sound alone · 0.80 corr

📖 About

VisTouch was constructed for Cross-Modal Semantic Communications (IEEE Wireless Communications, 2022) by controlling a robot arm (UR3 + RH56BF3 dexterous hand) to press and slide across everyday materials while a camera, a microphone, and a tactile force sensor record the same contact event simultaneously.

The full VisTouch research corpus spans 47 material categories and contains million-scale raw sensor observations across video frames, audio waveform samples, and tactile measurements. The companion paper reports 1000+ synchronized video–audio–haptic signal pairs over all 47 categories, supporting cross-modal semantic encoding, retrieval, and haptic signal recovery research.

This repository is the first curated public benchmark release: 2000 timestamp-aligned audio/tactile/video triplets from 8 representative material categories used in the paper's evaluation — brass · linen · paper · polyester · silk · spandex · stone · wood — with a predefined cross-pressure train/test split. Every released sample is a genuine sensor capture; no synthetic data is included. Future updates will progressively release more of the 47-category corpus, additional paths, views, and benchmark tasks.

Capture rig: dexterous hand + microphone + tactile sensor, fixed camera view.

📥 Download

The data files are hosted externally — this repository ships with an empty dataset/ folder:

SourceLink
🌐 Google Drivedrive.google.com/drive/folders/1U2qW1Oqbkj-...
☁️ Baidu Netdisk[pan.baidu.com/s/1W4cRtxgY9SL85HdnZhoFJA](https://pan.baidu.com/s/

项目描述

A large-scale synchronized vision–touch–audio dataset of robotic sliding contact

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

xiejunjie524
精选
xiejunjie524 GitHub avatar

handdraw-story-video

A tool that converts 7–9 hand-drawn story keyframes into a 35–45 second vertical video with progressive line art and color animation, configurable via JSON and renders with HyperFrames and GSAP.

设计与创意
692
bytec-ai
精选
bytec-ai GitHub avatar

book-video-factory

A portable multi-account short video production skill for books that manages accounts, scripts, images, voiceovers, subtitles, and final exports through a unified workspace.

设计与创意
136
cvg
精选
cvg GitHub avatar

vidmap

VidMap is an offline video-based Structure-from-Motion (SfM) system that combines temporal tracks, loop closures, metric depth, and global optimization to estimate camera poses, intrinsics, and a sparse 3D map.

设计与创意
64