Skip to content
VLA 線 · DEEP DIVE ARCHIVEVLA 线 · DEEP DIVE ARCHIVE

VLA 深度追蹤VLA 深度追踪

Vision-Language-Action:讓機器人看→想→做的端到端模型Vision-Language-Action:让机器人看→想→做的端到端模型

METHOD FAMILY TRENDS

data thru 2026-07-21 · 667 papers · 50d window · 15 families
high
▼ 3 declining 15 families · 243 papers covered
FAMILY MOM 7d 14d 30d Δ7d Δ14d Δ30d CHART ST
Lang. Grounding 66 109 205 1.79x 0.96x 0.96x
RL Fine-tuning 32 46 91 1.32x 0.86x 0.94x
World Model 28 51 127 1.15x 0.57x 1.31x
Tactile 19 32 54 0.78x 0.61x 0.56x
Multi-Task 18 35 64 0.74x 0.66x 0.66x
Human-Robot 17 33 66 0.70x 0.63x 0.68x
Long Horizon 15 26 57 0.62x 0.49x 0.59x
Flow Matching 13 25 47 0.53x 0.47x 0.48x
Dexterous Hand 10 24 42 0.41x 0.46x 0.43x
Diffusion Policy 7 13 17 0.29x 0.25x 0.17x
Cross-Embodiment 5 9 18 0.21x 0.17x 0.19x
Mobile Manip. 5 5 11 0.21x 0.09x 0.11x
Sim-to-Real 5 10 27 0.21x 0.19x 0.28x
3D Repr. 3 5 12 0.12x 0.09x 0.12x
Instr. Tuning 0 2 3 0.00x 0.04x 0.03x
COMPETITION PAIRS 6 matchups · hover for details
VLA vs WAM
Paradigm war: end-to-end action prediction vs world model planning
Lang. Ground. vs World Model
70%
30%
9.9% · x1.79 ratio 2.36 4.2% · x1.15
Lang. Grounding

Language Grounding: connecting natural language instructions to robot actions; vision-language-action alignment

World Model

World Model: learned environment simulator (Dreamer, UniSim); enables planning via imagination without real-world interaction

Why they compete

The central paradigm war in embodied AI. VLA (Vision-Language-Action) maps observations directly to actions end-to-end — simple, scalable, but needs massive data and generalizes poorly. WAM (World-Action Model) first learns how the world works, then plans actions through mental simulation — better generalization and data efficiency, but world models are often inaccurate. The boundary is blurring: Pi0.5 uses flow matching (generative, WAM-like), GR00T adds video prediction. The winner likely is a hybrid.

ACTION HEAD ROUTE
Continuous action generation: denoising vs optimal transport
Diffusion Pol. vs Flow Matching
35%
65%
1.1% · x0.29 ratio 0.54 1.9% · x0.53
Diffusion Policy

Diffusion Policy: iterative denoising process (DDPM) to generate continuous robot actions; strong on multi-modal action distributions

Flow Matching

Flow Matching: optimal-transport-based generative model (e.g. Pi0); faster inference than diffusion with comparable quality

Why they compete

Both generate continuous actions from the same VLA backbone but take different mathematical routes: diffusion iteratively denoises random noise into actions (slow, expressive), while flow matching uses optimal transport for a direct trajectory (fast, efficient). If flow matching matches diffusion quality, it could replace it as the default action head.

POST-TRAINING ROUTE
Model adaptation: supervised tuning vs reward optimization
Instr. Tuning vs RL Fine-tune
0%
100%
0.0% · x0.00 ratio 0.00 4.8% · x1.32
Instr. Tuning

Instruction Tuning: supervised fine-tuning (SFT) on language-action pairs; simpler but limited to offline data distribution

RL Fine-tuning

RL Fine-tuning: post-training with PPO/DPO/GRPO reward signals; enables online improvement beyond demonstration data

Why they compete

After pretraining a VLA, two competing strategies exist: SFT directly imitates expert demonstrations (simple, stable), while RL fine-tuning (GRPO/DPO) optimizes a reward signal to go beyond the demonstration distribution. RL can discover novel strategies but is harder to stabilize.

LEARNING SIGNAL
Training paradigm: imagination-based vs reward-based
World Model vs RL Fine-tune
47%
53%
4.2% · x1.15 ratio 0.88 4.8% · x1.32
World Model

World Model: learned environment simulator (Dreamer, UniSim); enables planning via imagination without real-world interaction

RL Fine-tuning

RL Fine-tuning: post-training with PPO/DPO/GRPO reward signals; enables online improvement beyond demonstration data

Why they compete

World models learn by predicting the future (imagination-based planning), while RL learns from reward feedback. If world models become accurate enough, they could reduce the need for expensive real-world RL exploration.

MANIPULATION SENSING
Manipulation approach: tactile feedback vs dexterous control
Tactile vs Dext. Hand
66%
34%
2.9% · x0.78 ratio 1.90 1.5% · x0.41
Tactile

Tactile Sensing: force/torque and GelSight contact sensors; provides direct manipulation feedback for delicate tasks

Dexterous Hand

Dexterous Hand: multi-finger manipulation control; achieves fine-grained object interaction without dedicated sensors

Why they compete

Two approaches to dexterous manipulation: tactile sensing adds explicit touch feedback (hardware cost, rich signal), while dexterous hand control relies on proprioception and vision alone (simpler hardware, harder control). The winner depends on sensor cost-to-performance ratio.

TRANSFER APPROACH
Domain bridging: simulation transfer vs cross-embodiment
Sim2Real vs Cross-Embod.
50%
50%
0.8% · x0.21 ratio 1.00 0.8% · x0.21
Sim-to-Real

Sim-to-Real: train in simulation, deploy on real hardware; uses domain randomization to bridge the reality gap

Cross-Embodiment

Cross-Embodiment: transfer policies across different robot morphologies; aims for universal robot foundation models

Why they compete

Sim-to-Real trains one robot in simulation then transfers (cheap data, reality gap risk), while Cross-Embodiment trains across multiple real robots directly (expensive data, natural generalization). The approaches represent different bets on where generalization should happen.

EMERGING SIGNALS

2026-07-21 · 34/191 unmatched · 7d window
3 signals
TERM COUNT AGE VELOCITY STATUS SAMPLE
closed loop
8 17d ~ x0.0 CANDIDATE RADAR: Closed-Loop Robotic Data Generation via Sem...
robot manipulation
6 13d -- x3.0 RISING Exo2EgoPose: Leveraging Exocentric Demonstrations ...
embodied ai
6 24d ~ x-1.0 CANDIDATE EgoExoMoCap: Distributed Ego-Exo Human Motion Capt...

TOP INSTITUTIONS

30d window · 20 labs tracked · VLA domain
20 active / 30d
INSTITUTION TOTAL BEST LAST SEEN ACTIVITY
1 LIBERO Team 8 🔧 07-17
2 Berkeley 7 07-15
3 Stanford 5 07-16
CMU 3 🔧 07-14
Fox 2 🔧 07-09
Zhu 2 🔧 07-14
Sadigh 2 🔧 07-16
北大 2 🔧 07-21
Wang 2 🔧 07-21
NVIDIA 3 07-09
Physical Intelligence 3 🔧 07-15
MIT 2 🔧 07-14
UW 1 🔧 06-30
Mishra 1 🔧 06-26
Malik 1 🔧 06-26
Li 1 🔧 06-26
UMich 1 🔧 07-01
Berenson 1 🔧 07-01
UCSD 1 🔧 07-01
Song 1 🔧 07-01

📐 理論文章庫📐 理论文章库

291 查看 GitHub 全庫查看 GitHub 全库
最近 2 週最近 2 周 25
6 天前 foundation

动作 QFormer:动作监督下的结构化表征塑造 (Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models)

在 GitHub 閱讀在 GitHub 阅读
6 天前 vla core

RoboTTT:通过测试时训练将 VLA 上下文扩展至 8K 时间步 (RoboTTT: Context Scaling for Robot Policies)

在 GitHub 閱讀在 GitHub 阅读
6 天前 vla core

迈向类人物理智能:面向机器人操作的终身视觉-语言-动作学习 (Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation)

在 GitHub 閱讀在 GitHub 阅读
7 天前 vla core

AeroAct: 动作中心世界-动作模型用于语言条件四旋翼飞行 (AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight)

在 GitHub 閱讀在 GitHub 阅读
7 天前 foundation

主动式真实世界因子评估框架 (Active Real-World Factor-Based Evaluation for Generalist Robot Policies)

在 GitHub 閱讀在 GitHub 阅读
8 天前 vla core

HELP:面向 VLA 后训练的人类高效流水线 (HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation)

在 GitHub 閱讀在 GitHub 阅读
8 天前 vla core

UniSteer:统一噪声引导的高效人类指导 VLA 自适应 (UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation)

在 GitHub 閱讀在 GitHub 阅读
8 天前 frontier

基于轨迹分割的人类高效大规模机器人后训练框架 (HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation)

在 GitHub 閱讀在 GitHub 阅读
9 天前 vla core

诊断 Agent 编排 VLA 技能组合中的语义交接失败 (Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition)

在 GitHub 閱讀在 GitHub 阅读
9 天前 tactile

在哪里触碰,如何接触:分层 RL-MPC 几何感知 Sim-to-Real 操作框架 (Where to Touch, How to Contact: A Hierarchical RL-MPC Framework for Geometry-Aware Sim-to-Real Manipulation)

在 GitHub 閱讀在 GitHub 阅读
9 天前 foundation

RoboWorld:面向通用机器人策略评估的快速可靠神经仿真器 (RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation)

在 GitHub 閱讀在 GitHub 阅读
10 天前 vla core

Harness VLA:通过记忆引导代理将冻结 VLA 转化为可靠操作原语 (Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents)

在 GitHub 閱讀在 GitHub 阅读
10 天前 vla core

DenseReward:通过失败合成实现密集奖励学习 (DenseReward: Dense Reward Learning via Failure Synthesis for Robotic Manipulation)

在 GitHub 閱讀在 GitHub 阅读
11 天前 vla core

混合帧策略:双臂移动操作的多帧动作去噪 (Mixture of Frames Policy: Multi-Frame Action Denoising for Bimanual Mobile Manipulation)

在 GitHub 閱讀在 GitHub 阅读
11 天前 foundation

Embodied-R1.5:通过具身基础模型进化物理智能 (Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models)

在 GitHub 閱讀在 GitHub 阅读
11 天前 foundation

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models (Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models)

在 GitHub 閱讀在 GitHub 阅读
11 天前 rl

从失败中学更多:VLA 后训练的事后强化学习 (Learning More from Less: Reinforcement Learning from Hindsight)

在 GitHub 閱讀在 GitHub 阅读
12 天前 vla core

RoboStream:在视觉语言模型中编织时空推理与记忆 (RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics)

在 GitHub 閱讀在 GitHub 阅读
13 天前 foundation

扩散策略长上下文训练与评估深度拆解 (Training and Evaluating Diffusion Policies with Long Context Lengths)

在 GitHub 閱讀在 GitHub 阅读
13 天前 vla core

SeFA-Policy:选择性流对齐视觉运动策略 (SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment)

在 GitHub 閱讀在 GitHub 阅读
13 天前 vla core

第一人称视频语言模型能否同时捕捉手部和物体中心线索?(Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?)

在 GitHub 閱讀在 GitHub 阅读
更多文章 · 全部在 GitHub更多文章 · 全部在 GitHub 266
🏛️ VLA Core  ·  144
V-VLAPS:价值引导的 VLA 规划 (Value-Guided Planning for Vision-Language-Action Models) 紧凑世界模型中的空间关系 Grounding:指令泄漏与无目标动力学修复 (Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix) X-Foresight:通过预测世界模型实现视觉-动作联合因果预测网络 (X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling) FailSafe:VLA 模型的失败推理与恢复系统 (FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models) Lift3D-VLA:将 VLA 模型提升至 3D 几何与动力学感知操作 (Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation) SEAM:动作块平滑执行 (Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies) 学习语义原子技能用于多任务机器人操作 (Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation) 传输差异作为 VLA 模型的可靠性信号 (Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models) 通过免训练注意力重校准恢复 VLA 模型的语言 grounding (Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration) 运动聚焦潜在动作实现跨具身 VLA 人类视频训练 (Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos) VLSA: 即插即用安全约束层的 VLA 模型 (VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer) 解耦视频生成世界模型 (DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation) 人即人形:从主客体人类视频中零样本学习人形机器人控制 (Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments) 通过锚定机器人关键点的序列规划 (Sequential Planning via Anchored Robotic Keypoints — SPARK) 行为提示策略:用单条演示作为操作任务的 Prompt (Behavior Prompting Policy: Demonstrations as Prompts for Manipulation) RouterVLA:将烟雾测试转化为异构 VLA 选择的监督信号 (RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection) Flow Matching 策略的 RL 精调:用 CFM 损失差分替代似然比 (Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models) TIDAL:时间交错扩散与动作循环实现高频 VLA 控制 (TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control) 无接触,无担忧:通过视觉和本体感估计灵巧操作中的接触力 (NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation) CoRDE:概念先验路由扩散专家 (Concept-Prior Routed Diffusion Experts for Structural Generalization in Robot Manipulation) Wh0: 用生成式世界模型合成第一人称手部操作数据 (Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data) VQActFlow:向量量化動作流的多任務機器人操控 (VQActFlow: Vector-Quantized Action Mode Steering for Multi-Task Robot Manipulation) 用接地潜在动作世界模型从异构演示中学习 (Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models) Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting (Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting) Mix-QVLA:任务证据感知的 VLA 混合精度量化 (Mix-QVLA: Task-Evidence-Aware Mixed-Precision Quantization of Vision-Language-Action Models) ImageWAM:世界动作模型真的需要视频生成吗?还是只需要图像编辑?(ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?) ENPIRE:真实世界中的具身智能体策略自改进 (ENPIRE: Agentic Robot Policy Self-Improvement in the Real World) 域自适应扩散策略 (Domain Adaptive Diffusion Policy) MemoryWAM:高效世界动作建模与持久记忆 (MemoryWAM: Efficient World Action Modeling with Persistent Memory) VLA 连常识都不知道了?测量视觉-语言-动作模型中的常识与世界知识保留 (Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models) 流匹配 VLA 的不确定性量化 (Uncertainty Quantification for Flow-Based Vision-Language-Action Models) GeneralVLA-2:几何感知重建与受控记忆驱动机器人规划 (GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning) 检索替代微调:测试时扩展 VLA 至新任务 (Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time) LaWAM:隐空间世界动作模型 (Latent World Action Models for Efficient Dynamics-Aware Robot Policies) AcceRL:面向 VLA 的分布式异步强化学习与世界模型框架 (AcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models) 力感知世界动作模型:闭环接触丰富操作 (FAWAM: Force-Aware World Action Models for Closed-Loop Contact-Rich Manipulation) EquiDexFlow:接触约束的 SE(3) 等变灵巧抓取生成流 (EquiDexFlow: Contact-Grounded SE(3)-Equivariant Dexterous Grasp Generative Flows) GAE: 用通用动作专家释放 VLM 的物理潜能 (Unleashing Physical Potential of VLM with Generalizable Action Expert) 利用共形预测从稀疏人类反馈中学习机器人安全 (Learning Robot Safety from Sparse Human Feedback using Conformal Prediction) 从数字到物理:数字智能体作为物理智能的自主教练 (From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence) 具身可解释性:因果理解驱动 VLA 泛化 (Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models) World Pilot:用世界-动作先验引导 VLA 决策 (World Pilot: Steering Vision-Language-Action Models with World-Action Priors) 离散时间高斯过程混合物在机器人策略学习中的不合理有效性 (The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning) MIND-V:分层世界模型与RL物理对齐的长程机械操作 (MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment) LIBERO-Occ:通过视点想象克服场景遮挡的 VLA 框架 (LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination) ORCHID:分层扩散策略的在线自训练共适应 (Online Self-Training for Co-Adaptation in Hierarchical Diffusion Policies) 统一对象中心世界模型与扩散策略:多阶段机器人任务的分层框架 (Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks) LARA:潜在动作表示对齐 (Latent Action Representation Alignment for Vision-Language-Action Models) AEGIS:物理 AI 的備份反射機制 (A Backup Reflex for Physical AI) 3DThinkVLA:通过3D思维引导协同训练赋予VLA模型隐式3D空间推理能力 (3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training) 势函数引导的 Flow Matching 用于 VLA 策略优化 (Potential-Guided Flow Matching for Vision-Language-Action Policy Improvement) TempoVLA:速度可控的 Vision-Language-Action 策略 (TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies) 潜入场景:通过焦点计划生成打破视觉语言决策中的感知瓶颈 (Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation) Dream.exe:视频生成模型能否"梦想"可执行的机器人操作?(Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?) OMP:单步均值流策略与方向对齐 (One-step MeanFlow Policy with Directional Alignment) LEGS:在高斯泼溅世界微调免遥操作 VLA 实现人形机器人全身操作 (LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World) 集合监督扩散策略:通过修正学习动作分块扩散 (Set-Supervised Diffusion Policy: Learning Action-Chunking Diffusion through Corrections) 好的具身奖励模型需要坏行为数据 (Good Embodied Reward Models Need Bad Behavior Data) AnySlot:零样本槽位级放置的目标条件 VLA (Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement) Hyper-DP3:频域感知的3D扩散策略轻量化重构 (Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control) 双流扩散世界模型增强 VLA (Dual-Stream Diffusion for World-Model Augmented VLA) DynaFLIP:三模态动力学引导的机器人感知重构 (DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation) GaussianDream:前馈式 3D 高斯世界模型赋能机器人操作 (GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation) VLA-Trace:通过表征与行为追踪诊断 VLA 模型 (VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing) 动态混合渐进式参数高效专家库用于终身机器人学习 (Dynamic Mixture of Progressive Parameter-Efficient Expert Library for Lifelong Robot Learning) 神经隐式动作场:从离散路点到连续函数 (Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models) LEIA:交互式架构材料的世界模型 (LEIA: Learned Environment for Interactive Architected Materials) CogVLA:认知对齐的视觉-语言-动作模型 (CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification) 能力与鲁棒性不可兼得:VLA 模型的信息论边界 (Capability and Robustness Cannot Both Be Free: An Information-Theoretic Bound for Vision-Language-Action Models) SOLE-R1:视频语言推理作为机器人在线强化学习的唯一奖励信号 (SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning) 弥合语义-动作鸿沟:面向高效 VLA 推理的视觉 Token 剪枝 (Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference) World-VLA-Loop:视频世界模型与 VLA 策略的闭环联合学习 (Closed-Loop Learning of Video World Model and VLA Policy) INSIGHT: 推理时序列自省生成人工辅助触发器 (INference-time Sequence Introspection for Generating Help Triggers in VLA Models) LIBERO-PRO: 超越死记硬背的 VLA 鲁棒与公平评估 (LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization) 动静解耦高效长视界 VLA (Static-Dynamic Disentanglement for Efficient Multi-Frame Vision-Language-Action Models) VGAS:价值引导的动作块选择用于少样本 VLA 适配 (Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation) USIM and U0:面向通用水下机器人的 VLA 数据集与模型 (USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots) LACY:基于视觉语言模型的双向语言-动作循环,实现机器人自我改进操作 (LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation) WorldKV:通过检索与压缩实现高效世界记忆 (Efficient World Memory with World Retrieval and Compression) 噪声空间归因与分块边界伪影控制 (Noise-Space Attribution and Control of Chunk-Boundary Artifact) DSSP:基于全历史编码的扩散状态空间策略 (Diffusion State Space Policy with Full-History Encoding) 仅凭本体感知实现灵巧手内操作:本体感知 Transformer (Learning Robust Dexterous In-Hand Manipulation from Joint Sensors with Proprioceptive Transformer) 手在环中:通过无缝手-臂干预改善 VLA 灵巧操作策略 (Hand-in-the-Loop: Improving VLA Policies for Dexterous Manipulation via Seamless Hand-Arm Intervention) SWEET:用图像编辑做稀疏世界模型 (SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution) 面向长寿机器人:通过强化微调实现 VLA 持续学习 (Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning) OxyGen:面向多任务并行的 VLA 统一 KV Cache 管理 (Unified KV Cache Management for VLA Inference under Multi-Task Parallelism) CLARE: VLA 持续学习通过适配器路由与扩展 (Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion) 离线语义引导的 VLA 策略高效蒸馏 (Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation) 学习结果发散之处:通过概率分块掩码加速 VLA RL 后训练 (Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking) 分块自适应缓存加速扩散策略 (Block-wise Adaptive Caching for Accelerating Diffusion Policy) D-VLA:高并发分布式异步强化学习框架 (D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models) 在相机帧中统一机器人动作 (Unify Robot Actions in Camera Frame) CoWorld-VLA:多专家世界模型中的思考式推理 (CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving) ALAM: 代数一致潜在动作模型 (Algebraically Consistent Latent Action Model for Vision-Language-Action Models) UniJEPA:统一连续与离散表征学习的机器人策略 (UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning) 行为模式发现:微调多模态生成策略时防止模式坍塌 (Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies) 从想象未来到可执行动作:混合潜在动作机器人操作 (From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation) 统一噪声引导:高效人类指导的 VLA 适配 (Unified Noise Steering for Efficient Human-Guided VLA Adaptation) 视觉预见VLA的测试时训练 (Test-Time Training for Visual Foresight Vision-Language-Action Models) PriorVLA:保留先验的 VLA 微调框架 (Prior-Preserving Adaptation for Vision-Language-Action Models) VEGA:视觉编码器接地对齐实现空间感知 VLA (VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models) ALAM:代数一致的潜在转移(Algebraically Consistent Latent Transitions for Vision-Language-Action Models) Hydra-DP3:面向视觉运动控制的3D扩散策略频域瘦身 (Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control) Sword: 风格鲁棒的世界模型模拟器 (Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training) TAIL-Safe:面向模仿学习策略的任务无关安全监控框架 AsyncVLA:异步流匹配视觉-语言-动作模型 (AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models) 持续演化 VLA 技能知识 (Continually Evolving Skill Knowledge in Vision Language Action Model) MobileEgo Anywhere:用消费级手机采集长视界第一人称数据 (MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware) 动作到动作流匹配 (Action-to-Action Flow Matching) LaST-R1: 通过自适应物理潜在推理强化机器人操作 (Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning) 显式物理可行性能否提升 VLA 学习?一项实证研究 (Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study) FingerViP: 指尖视觉感知灵巧操作 (FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception) MolmoAct2:面向真实世界部署的动作推理模型 (MolmoAct2: Action Reasoning Models for Real-world Deployment) STEP:时空一致性预测的热启动视觉运动策略 (Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction) 用文字和图像思考:长程机器人操作的交错视觉-语言推理轨迹 (Thinking in Text and Images: Interleaved Vision-Language Reasoning Traces for Long-Horizon Robot Manipulation) VLA 受限于训练但具备新指令泛化能力 (VLAs are Confined yet Capable of Generalizing to Novel Instructions) MotuBrain:面向机器人控制的世界-动作统一生成模型 (MotuBrain: An Advanced World Action Model for Robot Control) PRTS:基于对比表示的基元推理与任务系统 (PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations) RopeDreamer:基于运动学递归状态空间模型的柔性线性物体动力学预测 (RopeDreamer: A Kinematic Recurrent State Space Model for Dynamics of Flexible Deformable Linear Objects) 弹性视觉智能体的架构模式语言 (A Pattern Language for Resilient Visual Agents) 从动作标签到动作集合:重新思考纠正反馈下的模仿学习动作监督 (From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback) 提升具身世界模型用于规划与控制 (Lifting Embodied World Models for Planning and Control) 将世界模型想象力蒸馏进 VLM:面向动态空间推理的训练框架 (World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning) RISE: 用组合世界模型实现 VLA 策略自改进 (Self-Improving Robot Policy with Compositional World Model) DIAL:通过潜在世界建模解耦意图与动作的端到端 VLA (DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA) HANDFUL:资源感知的序列灵巧操作 (Sequential Grasp-Conditioned Dexterous Manipulation with Resource Awareness) KERV:运动学校正推测解码用于具身 VLA 模型 (Kinematic-Rectified Speculative Decoding for Embodied VLA Models) RoboECC: VLA 模型的多因素感知云边协同部署框架 (RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models) SARM:阶段感知奖励建模用于长视界机器人操作 (SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation) 通过对象中心与几何接地提升杂乱环境下的 VLA 鲁棒性 (Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding) dWorldEval:基于离散扩散世界模型的-scalable 机器人策略评估 (Scalable Robotic Policy Evaluation via Discrete Diffusion World Model) GazeVLA:用人类注视学习操作意图 (Learning Human Intention for Robotic Manipulation) 行为克隆策略有多脆弱?通用对抗扰动攻击现代BC策略 (How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies) VistaBot: 视角鲁棒机器人操作通过时空感知视图合成 (View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis) 长视界操作:轨迹条件化 VLA 规划 (Long-Horizon Manipulation via Trace-Conditioned VLA Planning) 从噪声到意图:用残差桥接锚定生成式 VLA 策略 (From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges) PhysMem: 测试时物理记忆扩展 (Scaling Test-time Physical Memory for Robot Manipulation) 基于漂移的策略优化:面向在线机器人控制的单步原生策略学习 (Drift-Based Policy Optimization: Native One-Step Policy Learning for Online Robot Control) 世界-价值-动作模型:VLA 系统的隐式规划 (World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems) DeepThinkVLA:增强视觉-语言-动作模型的推理能力 (DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models) 长程记忆赋能 VLA 智能体在开放世界任务执行 (Long-Term Memory for VLA-based Agents in Open-World Task Execution) 从看到仿真:用数字表亲生成高保真仿真环境 (From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation) 力场流匹配:从单演示生成力觉数据学习 3D 顺应性策略 (Flow with the Force Field: Learning 3D Compliant Flow Matching Policies from Force and Demonstration-Guided Simulation Data) 无需微调部署 VLA:即插即用推理时策略引导 (Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion)
🏗️ Foundation & Training  ·  47
HiMoE-VLA:分层混合专家通用视觉-语言-动作策略 (Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies) StereoVLA:用立体视觉增强 VLA 的空间感知 (Enhancing Vision-Language-Action Models with Stereo Vision) LA4VLA:看不见也能行动——通过语言-动作预训练解耦VLA中的视觉依赖 (Learning to Act without Seeing via Language-Action Pretraining) WAM 自我回放:用世界模型生成伪轨迹实现持续模仿学习 (World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays) 先学会玩,再学会装配:灵巧手 Play Pretraining 的关键设计因素 (Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?) PAIWorld:多视图3D一致性世界基础模型 (PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation) 拿我的杯子!用视觉注意力提示个性化 VLA 模型 (Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting) DeMaVLA:面向可变形物体操作的通用 VLA 基础模型 (DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation) WEAVER:更好的、更快的、更长的——面向机器人操作的高效世界模型 Mana:铰接工具的灵巧操作 (Dexterous Manipulation of Articulated Tools) MotionWAM:迈向实时人形机器人世界动作模型 (MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation) Instant-Fold:单演示驱动的柔性物体折叠学习 (Instant-Fold: In-Context Imitation Learning for Deformable Object Manipulation) 部署中学习:面向通用机器人策略的车队规模强化学习 (Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies) SimuScene:从单图重建仿真就绪的组合 3D 场景 (SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image) Dexterity-BEV: 对齐3D世界与动作以增强策略泛化 (Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning) CrossVLA: 跨范式后训练与推理优化 (Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models) 通过保守 SFT 保护流匹配 VLA 的基础能力 (Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT) HEX:跨具身全身操控的类人对齐专家架构 (HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation) 通过模仿生成视频实现机器人操作(Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations) RoboEval:机器人操作的结构化与可扩展评估 (RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation) CLAMP: 3D 多视图对比预训练用于机器人操作 (Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining) GeCO:时间无条件流匹配用于自适应鲁棒机器人控制 (Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control) 基于基础模型先验的强化学习:让具身智能体自主高效学习 (Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own) 掩码世界模型:预测什么对机器人策略学习最重要 (Mask World Model: Predicting What Matters for Robust Robot Policy Learning) 人类数据是伪装成另一种形式的机器人数据:Danfei Xu 深度访谈(2026) 潜空间综述:语言模型的"原生思维空间"与具身智能的统一接口 免微调部署 VLA:即插即用推理时策略引导 (Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion) VLA 数据工程指南:从采集到训练的完整链路 StarVLA-α:简化视觉 - 语言 - 动作系统的强基线 (StarVLA-α: Reducing Complexity in Vision-Language-Action Systems) HY-Embodied-0.5:具身基础模型实战解析 (HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents) 联合训练 (Co-training) 数据处理 (Data Processing) 具身智能深度:数据飞轮与跨模态迁移 (Data Flywheel & Cross-modal Transfer) DCP:凸性检测规则与 CVX/CVXPY 建模心法 (Disciplined Convex Programming) DoRA:权重分解的低秩适配 (DoRA: Weight-Decomposed Low-Rank Adaptation) 评估体系详解 (Evaluation Protocols Deep Dive) Flash Attention: 高效 Transformer 推理的关键 🏗️ 基础理论 — ML 工具箱主线总纲 更新成本摊销:Doc-to-LoRA / Text-to-LoRA 让 LLM “瞬时内化” (Cost Amortization for Instant LLM Updates) 知识蒸馏 (Knowledge Distillation) Knowledge Insulation: 防止灾难性遗忘 当我们谈论 AI 推理的 KV Cache,我们在说什么? (KV Cache in LLM Inference) 终身模仿学习与多模态潜在回放 (Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment) VLA 文献核心技术归纳 (Literature Technical Review) VLA 数学必备:从直觉到实作 弹性模组化架构 Table 生成器(VLA Modular Pipelines) NeurIPS 2025 最佳论文:具身智能视角解读
🔧 Deployment & Hardware  ·  23
Mana:铰接工具的灵巧操作 (Dexterous Manipulation of Articulated Tools) DockAnywhere: 通过演示生成提升移动操作数据效率 (DockAnywhere: Data-Efficient Visuomotor Policy Learning for Mobile Manipulation via Novel Demonstration Generation) BLaDA:在 3DGS 场中桥接语言与功能性灵巧动作 (BLaDA: Bridging Language to Functional Dexterous Actions within 3DGS Fields) 迭代组合式数据生成用于机器人控制 (Iterative Compositional Data Generation for Robot Control) 🔧 部署与硬件 — 实战落地主线总纲 DexGrasp-Zero:形态对齐的零样本跨本体灵巧抓取策略 (DexGrasp-Zero: A Morphology-Aligned Policy for Zero-Shot Cross-Embodiment Dexterous Grasping) 中金人机系列05(灵巧手)→ VLA/控制/硬件的“可计算约束”框架(理论侧整理) 灵巧手机械学深度解析 (Dexterous Hand Mechanics) — 修订整合版 v2 机器人开可乐/发牌有多难?灵巧手:硬件路线 × 接触数学 × 数据金字塔(访谈摘录整理) EquiBim:双臂操作中的对称等变策略学习 (EquiBim: Learning Symmetry-Equivariant Policy for Bimanual Manipulation) GR-Dexter(ByteDance Seed):把 VLA 扩展到高自由度灵巧手的“硬件-数据-模型”全栈框架 抓取算法与仿真平台 (Grasp Algorithms & Simulation Platforms) House of Dextra: 灵巧手机器人形态 - 控制协同设计 (House of Dextra: Cross-embodied Co-design for Dexterous Hands) 产业视角:通用性与“元学习”路径(从一张路线图说起) Isaac Lab: GPU 加速的多模态机器人学习仿真框架 Lightning Grasp:Contact Field 驱动的超高速灵巧手抓取合成 (Lightning Grasp: Procedural Grasp Synthesis with Contact Fields) NVIDIA 的 AI 五层蛋糕:从能源到机器人应用的基础设施观 (AI Is a 5-Layer Cake) 英伟达物理 AI 的第一刀:为什么先砍向汽车 (Why NVIDIA's First Physical AI Wedge Hits Cars First) Physical Intelligence Layer:机器人基础模型 API 的产品化范式 (The Physical Intelligence Layer) RoboPocket:把“机器人博士”装进口袋的无本体即时策略迭代 (RoboPocket: Improve Robot Policies Instantly with Your Phone) 机械臂运动学、动力学与控制 (Robot Arm Kinematics, Dynamics & Control) 机器人动力学系统分类 (Classification of Robot Dynamical Systems) 机器人“开源基建”三分法:成果展示 / 生态绑定 / 基础设施(以 RoboParty Roboto_Origin 为例)
🌊 Diffusion & Flow  ·  15
VLA 的 RL 精调突破:Flow Policy Optimization (FPO) X-Diffusion: 跨具身人类演示训练扩散策略 (X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations) SnapFlow:流匹配 VLA 的单步动作生成 (SnapFlow: One-Step Action Generation for Flow-Matching VLAs via Progressive Self-Distillation) 动作生成范式详解 (Action Representations & Generation) 瓶颈定位:VLA 模型动作生成的边缘架构困境 (Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures) 闭环动作块动态校正:训练-free 扩散策略实时适应 (Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy) CoMo:从互联网视频学习连续潜在运动 (CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning) 收缩扩散策略:通过收缩微分方程实现鲁棒动作扩散 (Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations) 🌊 扩散与流匹配 — 动作生成主线总纲 扩散策略详解 (Diffusion Policy) Pi0 (π0) 代码解构:Flow Matching for VLA Pixel Motion Diffusion is What We Need for Robot Control (DAWN) 压缩鸿沟:为何离散 Tokenization 限制 VLA 模型Scaling (The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling) 传统动作生成方法 (Traditional Action Generation) 全模态共享 Token 空间:以 MM-ACT 为例的 VLA 进化论
🤚 Tactile Perception  ·  13
安全感知视触觉基准:面向可变形物体的物理约束机器人操作 (SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects) 唤醒触觉!MLLM 中的掩码隔离触觉对齐学习 (Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs) Event-VLA:动作条件化事件融合实现光照鲁棒 VLA (Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model) Mana:铰接工具的灵巧操作 (Dexterous Manipulation of Articulated Tools) Mana: Dexterous Manipulation of Articulated Tools 学会感知未来:DreamTacVLA 用于接触丰富操作 (Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation) 从抓取到插入:亚毫米级触觉增强精密装配 (From Reach to Insert: Tactile-Augmented Precision Assembly under Sub-Millimeter Tolerances) 语义接触场:类别级通用触觉工具操作 (Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation) TouchGuide: 推理时触觉引导视觉运动策略 (TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance) OmniUMI: 面向物理具身机器人学习的多模态交互接口 (Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction) ETac:轻量高效的触觉仿真框架,赋能灵巧操作学习 (ETac: A Lightweight and Efficient Tactile Simulation Framework for Learning Dexterous Manipulation) 多模态操作 via 多模态策略共识 (Multi-Modal Manipulation via Multi-Modal Policy Consensus) TAMEn: 触觉感知操作引擎用于接触丰富任务中的闭环数据收集 (TAMEn: Tactile-Aware Manipulation Engine for Closed-Loop Data Collection in Contact-Rich Tasks)
🔬 Frontier Research  ·  11
ENPIRE:真实世界中的具身智能体策略自进化 (ENPIRE: Agentic Robot Policy Self-Improvement in the Real World) GAE:释放 VLM 的物理潜能,以通用动作专家解耦推理与执行 (Unleashing Physical Potential of VLM with Generalizable Action Expert) cuRoboV2:高自由度机器人的动力学感知运动生成 (cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots) 人工三元智能:生物启发的物理 AI 传感器优先架构 (Artificial Tripartite Intelligence: A Bio-Inspired, Sensor-First Architecture for Physical AI) IGen: 从开放世界图像可扩展生成机器人学习数据 (IGen: Scalable Data Generation for Robot Learning from Open-World Images) StaMo:从紧凑状态表示中涌现通用机器人运动 (StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation) StaMo:从紧凑状态表示中涌现通用机器人运动 (StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation) 用自由语言指令操控人形机器人:统一运动词汇的大型语言动作模型 (Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary) Déjà Vu:具身智能的经验反馈学习框架 (Dejavu: Towards Experience Feedback Learning for Embodied Intelligence) 你有一张金票:用单个噪声向量提升生成式机器人策略 (You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector) RoSHI: 野外便携式全身动捕套装 (RoSHI: A Versatile Robot-oriented Suit for Human Data In-the-Wild)

🏆 SOTA 排行SOTA 排行

Evo-SOTA 完整榜Evo-SOTA 完整榜 30
CALVIN ABCD-D 飽和饱和 avg_len
# Model Score vs Prev Date Paper
1 Xiaomi-Robotics-0 4.8 Flower VLA +0.13 2026-07-17 arxiv →
2 Xiaomi-Robotics-0 4.8 Flower VLA +0.13 2026-07-10 arxiv →
3 MMaDA-VLA 4.78 Xiaomi-Robotics-0 +0.03 2026-07-17 arxiv →
4 MMaDA-VLA 4.78 Xiaomi-Robotics-0 +0.03 2026-07-10 arxiv →
5 NIAF 4.66 GR-2 +0.02 2026-07-17 arxiv →
6 NIAF 4.66 GR-2 +0.02 2026-07-10 arxiv →
7 AVA-VLA 4.65 NIAF +0.18 2026-07-17 arxiv →
8 AVA-VLA 4.65 NIAF +0.18 2026-07-10 arxiv →
9 NS-VLA 4.56 AtomicVLA +0.29 2026-07-17 arxiv →
10 NS-VLA 4.56 AtomicVLA +0.29 2026-07-10 arxiv →
11 Flower VLA 4.35 RoboUniview +0.49 2026-07-17 arxiv →
12 Flower VLA 4.35 RoboUniview +0.49 2026-07-10 arxiv →
13 MCIL 1.82 2026-07-17 arxiv →
14 MCIL 1.82 2026-07-10 arxiv →
LIBERO standard-opensource 飽和饱和 average
# Model Score vs Prev Date Paper
1 LaST-R1 99.8 Abot-M0.5 +0.40 2026-07-17 arxiv →
2 PLD 99.17 NS-VLA +0.57 2026-07-17 arxiv →
3 PLD 99.17 NS-VLA +0.57 2026-07-10 arxiv →
4 PriorVLA 99.1 GeoAlign +0.10 2026-07-17 arxiv →
5 PriorVLA 99.1 GeoAlign +0.10 2026-07-10 arxiv →
LIBERO Plus standard-closed total
# Model Score vs Prev Date Paper
1 GEAR-VLA 88.7 TAG +1.46 2026-07-17 arxiv →
2 ACoT-VLA 86.6 pi0.5 +0.90 2026-07-17 arxiv →
3 CorridorVLA 83.21 NS-VLA +3.81 2026-07-17 arxiv →
MetaWorld standard-opensource average
# Model Score vs Prev Date Paper
1 FabriVLA 90 LA4VLA-1B +2.47 2026-07-17 arxiv →
2 FabriVLA 90 LA4VLA-1B +2.47 2026-07-12 arxiv →
3 MPI 86 iRe-VLA +3.00 2026-07-17 arxiv →
4 ALAM 85 OneWM-VLA +23.72 2026-07-17 arxiv →
RoboCasa-GR1-Tabletop standard-opensource avg_success_rate
# Model Score vs Prev Date Paper
1 WALA 75.2 DIAL +5.00 2026-07-17 arxiv →
2 PhysBrain 1.0 64.5 JoyAI-RA 0.1 +1.30 2026-07-17 arxiv →
RoboChallenge standard-opensource score
# Model Score vs Prev Date Paper
1 DM0 72.25 Giga-Brain-0.1 +3.91 2026-07-17 arxiv →
2 StarVLA-alpha 54.5 2026-07-17 arxiv →