ASD-FEAT: Multi-Modal Infant Video Dataset & Benchmark
ASD-FEAT (ASD – Feature Extraction And Tracking) is a longitudinal, multi-modal dataset derived from video recordings of infant–adult play sessions, released with the full training and evaluation code. It supports reproducible machine-learning research on early autism spectrum disorder (ASD) risk prediction.
What’s released
- Dataset (Zenodo, DOI 10.5281/zenodo.22261227): derived features only — face and eye landmarks, head and gaze pose, facial action units, I3D video embeddings, and mel spectrograms — with frame-level expert behavior labels and de-identified metadata. No raw video or audio is distributed.
- Code (github.com/sidrah-liaqat/asd-feat, MIT license): the two-stage pipeline that reproduces the paper’s results.
The pipeline
- Behavior detection: a multi-modal transformer encoder fuses the per-frame feature streams to detect four child behaviors (Look-Face, Look-Object, Smile, Social-Vocalization) at the frame level.
- Risk prediction: frame-level detections are aggregated into per-session features that feed an ASD-risk classifier.
The fully automated pipeline reaches 76.2% accuracy and an AUROC of 0.82 for ASD classification. Classifiers trained on manually coded behaviors reach 81.3% and 0.88.
Paper: S. Liaqat et al., “ASD-FEAT: A Multi-Modal Infant Video-Derived Dataset for Early ASD Risk Prediction,” under review at Data-centric Machine Learning Research (DMLR).
Code on GitHub →Dataset — DOI 10.5281/zenodo.22261227 →Code DOI 10.5281/zenodo.22309762 →