Articles tagged
multimodal
2 articles
Orca: What If Next-Token, Next-Frame, and Next-Action Are the Same Task?
Orca (BAAI) replaces next-token, next-frame, and next-action prediction with a single Next-State-Prediction objective. A frozen 4B backbone pre-trained on 12.5K hours of video — with zero action labels — feeds three lightweight readouts, and the action readout, trained on just 200 trajectories per task, beats π0.5 on OOD robot manipulation (32.4 vs 29.4).
world-modelsnext-state-predictionrepresentation-learning
HeLo – A New Path for Multimodal Emotion Recognition
HeLo predicts a distribution over emotions instead of one label, fusing ECG and video via cross-attention, entropic optimal transport and a label-correlation loss.
multimodalemotionmachine learning