VisTouch is a public benchmark release of 2000 timestamp-aligned audio/tactile/video triplets from 8 material classes, captured by a robotic sliding-contact rig, with baselines for material recognition, tactile super-resolution, cross-modal retrieval and generation.

Stars

34

7-day growth

No data

Forks

3

Open issues

0

License

NOASSERTION

Last updated

2026-07-31

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It is one of the first large-scale synchronized vision-touch-audio datasets for robotic sliding contact, offering real sensor captures, a cross-pressure generalization split, and reproducible lightweight baselines.

Who it is for

  • Robotic touch and tactile sensing researchers
  • Multimodal machine learning researchers
  • Cross-modal semantic communications community
  • Material recognition and haptic signal processing developers

Use cases

  • Training and benchmarking multimodal material recognition from audio, tactile, and video
  • Tactile force signal super-resolution from low-rate or noisy input
  • Cross-modal retrieval between touch and sound
  • Generating haptic/tactile force curves from audio-only input

Strengths

  • 2000 genuine synchronized audio/tactile/video triplets across 8 materials with timestamp alignment
  • Benchmarks show clear gains over chance/naive baselines: 76.0% material accuracy, tactile SR 0.996 correlation
  • Predefined train/test split by contact force tests cross-pressure generalization
  • Code and data licenses are permissive (MIT/CC-BY-4.0) and baselines are CPU-trainable

Considerations

  • Only 8 of 47 material categories are in the current public release
  • Dataset files must be downloaded externally; repository ships with empty dataset folder
  • Baselines are intentionally lightweight and not meant as state-of-the-art results

README quick start

A large-scale synchronized vision–touch–audio dataset of robotic sliding contact


🎬 What you can do with VisTouch

  🏷️ Material Recognition
  fused audio+tactile+video · 76.0% test acc (chance 12.5%)


  
  🧵 Tactile Super-Resolution
  clean 100Hz force curve from noisy low-rate input · −85% MSE, 0.996 corr




  
  🔍 Cross-Modal Retrieval
  touch query → matching sound clip · 2× over chance Recall@1/@5


  
  🔄 Cross-Modal Generation
  haptic recovery: tactile force curve generated from sound alone · 0.80 corr

📖 About

VisTouch was constructed for Cross-Modal Semantic Communications (IEEE Wireless Communications, 2022) by controlling a robot arm (UR3 + RH56BF3 dexterous hand) to press and slide across everyday materials while a camera, a microphone, and a tactile force sensor record the same contact event simultaneously.

The full VisTouch research corpus spans 47 material categories and contains million-scale raw sensor observations across video frames, audio waveform samples, and tactile measurements. The companion paper reports 1000+ synchronized video–audio–haptic signal pairs over all 47 categories, supporting cross-modal semantic encoding, retrieval, and haptic signal recovery research.

This repository is the first curated public benchmark release: 2000 timestamp-aligned audio/tactile/video triplets from 8 representative material categories used in the paper's evaluation — brass · linen · paper · polyester · silk · spandex · stone · wood — with a predefined cross-pressure train/test split. Every released sample is a genuine sensor capture; no synthetic data is included. Future updates will progressively release more of the 47-category corpus, additional paths, views, and benchmark tasks.

Capture rig: dexterous hand + microphone + tactile sensor, fixed camera view.

📥 Download

The data files are hosted externally — this repository ships with an empty dataset/ folder:

SourceLink
🌐 Google Drivedrive.google.com/drive/folders/1U2qW1Oqbkj-...
☁️ Baidu Netdisk[pan.baidu.com/s/1W4cRtxgY9SL85HdnZhoFJA](https://pan.baidu.com/s/

Description

A large-scale synchronized vision–touch–audio dataset of robotic sliding contact

Related repositories

Similar projects matched by category, topics, and programming language.

xiejunjie524
Featured
xiejunjie524 GitHub avatar

handdraw-story-video

A tool that converts 7–9 hand-drawn story keyframes into a 35–45 second vertical video with progressive line art and color animation, configurable via JSON and renders with HyperFrames and GSAP.

Design & Creative
692
bytec-ai
Featured
bytec-ai GitHub avatar

book-video-factory

A portable multi-account short video production skill for books that manages accounts, scripts, images, voiceovers, subtitles, and final exports through a unified workspace.

Design & Creative
136
cvg
Featured
cvg GitHub avatar

vidmap

VidMap is an offline video-based Structure-from-Motion (SfM) system that combines temporal tracks, loop closures, metric depth, and global optimization to estimate camera poses, intrinsics, and a sparse 3D map.

Design & Creative
64