Early detection of Autism Spectrum Disorder (ASD) is crucial for timely intervention, yet current diagnostic methods rely heavily on subjective behavioral observations. Skeletal motion analysis offers a promising alternative, as people with ASD often exhibit distinctive movement patterns that can be captured through motion tracking systems. However, it remains unclear which modeling strategies best exploit the temporal structure of motion sequences and whether combining learned representations with domain-informed features improves classification. In this work, we investigate sequence-based neural architectures for ASD detection from motion capture data. We compare recurrent and attention-based models, evaluate temporal downsampling strategies for handling sequences of varying lengths, and explore fusion mechanisms that integrate deep features with hand-crafted motion descriptors. Through ablation studies, we further explore which body regions carry the most discriminative information for classification. Extensive experiments on two publicly available datasets with contrasting sequence characteristics demonstrate that our approach achieves over 96% accuracy, substantially outperforming traditional machine learning methods. Our findings highlight the effectiveness of feature fusion and aim to guide the design of motion-based ASD screening systems.
Temporal Modeling with Feature Fusion for Autism Spectrum Disorder Detection from Skeletal Motion
La Quatra M.;Cammarata V.;Conti V.;Salerno V. M.;Sorce S.;Cilia N. D.
2027-01-01
Abstract
Early detection of Autism Spectrum Disorder (ASD) is crucial for timely intervention, yet current diagnostic methods rely heavily on subjective behavioral observations. Skeletal motion analysis offers a promising alternative, as people with ASD often exhibit distinctive movement patterns that can be captured through motion tracking systems. However, it remains unclear which modeling strategies best exploit the temporal structure of motion sequences and whether combining learned representations with domain-informed features improves classification. In this work, we investigate sequence-based neural architectures for ASD detection from motion capture data. We compare recurrent and attention-based models, evaluate temporal downsampling strategies for handling sequences of varying lengths, and explore fusion mechanisms that integrate deep features with hand-crafted motion descriptors. Through ablation studies, we further explore which body regions carry the most discriminative information for classification. Extensive experiments on two publicly available datasets with contrasting sequence characteristics demonstrate that our approach achieves over 96% accuracy, substantially outperforming traditional machine learning methods. Our findings highlight the effectiveness of feature fusion and aim to guide the design of motion-based ASD screening systems.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


