Architecture-data matching for EEG-EMG decoding: compact deep models match classical spectral decoders on the WAY-EEG-GAL grasp-and-lift dataset

Modern deep learning has broadened the tools available for non-invasive neural decoding, but its advantage over well-engineered classical pipelines remains unclear at clinical neural-engineering sample sizes. We compared four classical decoders, three multilayer perceptron (MLP) variants, and three deep-learning architectures (an EEGNet-style compact convolutional network, a four-layer Transformer encoder trained from scratch, and a graph attention network) on the public WAY-EEG-GAL grasp-and-lift dataset (12 participants, 3,528 trials). Models were evaluated using leave-one-subject-out (LOSO) cross-validation to decode object weight (165, 330, and 660 g) and grasp-surface friction (sandpaper, suede, and silk). After Benjamini–Hochberg false discovery rate (BH-FDR) correction within the primary/robustness family, none of the deep-versus-best-classical comparisons reached significance. The graph attention network led nominally on weight (0.643), and the compact convolutional network (CNN) led nominally on surface (0.565), but neither exceeded the best classical baselines (HGBM = 0.617 for weight; logistic regression = 0.562 for surface). Two-direction bandwidth controls showed that the nominal graph neural network (GNN) advantage on weight reflected access to higher-frequency electromyography (EMG) content rather than a robust architectural gain. The GNN weight signal was concentrated in the first 500 ms of sustained hold (0.639 early vs. 0.540 late, q = 0.007), whereas the compact CNN was comparatively robust to bandwidth and phase. A four-layer Transformer trained from scratch underperformed (q = 0.003 for both tasks), consistent with a parameter–data mismatch at n = 12. Modality ablation showed that both tasks were dominated by peripheral EMG features: electroencephalography (EEG)-only decoding was near chance for weight (0.34–0.37) and surface (0.36–0.38; all q < 0.0015 vs. EMG-only). EMG-only HGBM achieved the highest accuracy in the study (0.712 for weight and 0.584 for surface), and adding EEG channels reduced HGBM weight decoding (q = 0.007) and CNN surface decoding (q = 0.012). In contrast, the graph attention network was robust to EEG-induced fusion dilution (q > 0.5), consistent with attention-based down-weighting of low-information channels. Together, bandwidth, modality, phase, and conditioning controls indicate that, at clinical sample sizes, the critical design choice is not architectural complexity but modality-relevant channel selection and the ability to suppress low-information channels.