handdraw-story-video
A tool that converts 7–9 hand-drawn story keyframes into a 35–45 second vertical video with progressive line art and color animation, configurable via JSON and renders with HyperFrames and GSAP.
它提供了大规模同步视觉-触觉-音频的真实传感器数据,并设计了跨压力泛化划分与可复现的轻量基线,是多模态触觉感知研究的重要公开资源。
A large-scale synchronized vision–touch–audio dataset of robotic sliding contact
🏷️ Material Recognition
fused audio+tactile+video · 76.0% test acc (chance 12.5%)
🧵 Tactile Super-Resolution
clean 100Hz force curve from noisy low-rate input · −85% MSE, 0.996 corr
🔍 Cross-Modal Retrieval
touch query → matching sound clip · 2× over chance Recall@1/@5
🔄 Cross-Modal Generation
haptic recovery: tactile force curve generated from sound alone · 0.80 corr
VisTouch was constructed for Cross-Modal Semantic Communications (IEEE Wireless Communications, 2022) by controlling a robot arm (UR3 + RH56BF3 dexterous hand) to press and slide across everyday materials while a camera, a microphone, and a tactile force sensor record the same contact event simultaneously.
The full VisTouch research corpus spans 47 material categories and contains million-scale raw sensor observations across video frames, audio waveform samples, and tactile measurements. The companion paper reports 1000+ synchronized video–audio–haptic signal pairs over all 47 categories, supporting cross-modal semantic encoding, retrieval, and haptic signal recovery research.
This repository is the first curated public benchmark release: 2000 timestamp-aligned audio/tactile/video triplets from 8 representative material categories used in the paper's evaluation — brass · linen · paper · polyester · silk · spandex · stone · wood — with a predefined cross-pressure train/test split. Every released sample is a genuine sensor capture; no synthetic data is included. Future updates will progressively release more of the 47-category corpus, additional paths, views, and benchmark tasks.
Capture rig: dexterous hand + microphone + tactile sensor, fixed camera view.
The data files are hosted externally — this repository ships with an empty
dataset/ folder:
| Source | Link |
|---|---|
| 🌐 Google Drive | drive.google.com/drive/folders/1U2qW1Oqbkj-... |
| ☁️ Baidu Netdisk | [pan.baidu.com/s/1W4cRtxgY9SL85HdnZhoFJA](https://pan.baidu.com/s/ |
A large-scale synchronized vision–touch–audio dataset of robotic sliding contact
根据分类、Topic 和编程语言匹配的相似项目。
A tool that converts 7–9 hand-drawn story keyframes into a 35–45 second vertical video with progressive line art and color animation, configurable via JSON and renders with HyperFrames and GSAP.
A portable multi-account short video production skill for books that manages accounts, scripts, images, voiceovers, subtitles, and final exports through a unified workspace.
VidMap is an offline video-based Structure-from-Motion (SfM) system that combines temporal tracks, loop closures, metric depth, and global optimization to estimate camera poses, intrinsics, and a sparse 3D map.