Skip to content

VLA 線 · Vision-Language-Action

VLA 研究日報VLA 研究日报

VISION-LANGUAGE-ACTION · cs.RO + cs.AI + cs.LG


Vision-Language-Action(VLA)機器人系統 — 整合 cs.RO、cs.AI、cs.LG 三條 arxiv 流。 重點追蹤 flow matching、世界模型、具身推理等前沿方向,每日 09:00 CST 由 Qwen3.5-Plus 自動評級Vision-Language-Action(VLA)机器人系统 — 整合 cs.RO、cs.AI、cs.LG 三条 arxiv 流。 重点追踪 flow matching、世界模型、具身推理等前沿方向,每日 09:00 CST 由 Qwen3.5-Plus 自动评级

— 2026 年 7 月 —
5 天前
⚡ 1 🔧 4 📖 17

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

22 篇
⚡ 1 🔧 9 📖 17

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

27 篇
🔧 15 📖 17

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

32 篇
🔧 9 📖 17

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

26 篇
🔧 10 📖 19

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

29 篇
🔧 10 📖 15

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

25 篇
⚡ 1 🔧 11 📖 18

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

30 篇
⚡ 1 🔧 11 📖 18

GeoProp: Grounding Robot State in Vision for Generalist Manipulation

30 篇
🔧 9 📖 23

Learning 4D Geometric Priors for Inference-Efficient World Action Models

32 篇
🔧 8 📖 20

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

28 篇
📖 1

High-resolution real-time mechanochromic tactile sensors

1 篇
📖 1

High-resolution real-time mechanochromic tactile sensors

1 篇
🔧 11 📖 15

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

26 篇
⚡ 1 🔧 11 📖 22

EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories

34 篇
🔧 16 📖 12

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning

28 篇
🔧 14 📖 15

Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning

29 篇
— 2026 年 6 月 —
🔧 10 📖 13

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience

23 篇
📖 1

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

1 篇
⚡ 1 🔧 16 📖 15

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

32 篇
🔧 15 📖 15

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control

30 篇
⚡ 1 🔧 13 📖 22

Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization

36 篇
⚡ 1 🔧 15 📖 12

MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation

28 篇
⚡ 1 🔧 13 📖 14

WorkBenchMark: A LEGO-Based Assembly Benchmark with an Assembly-by-Disassembly Baseline for the Smart Manufacturing League

28 篇
⚡ 1 🔧 22 📖 6

Guava: An Effective and Universal Harness for Embodied Manipulation

29 篇
🔧 17 📖 11

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

28 篇
🔧 15 📖 13

$\mu_0$: A Scalable 3D Interaction-Trace World Model

28 篇
🔧 14 📖 16

$\mu_0$: A Scalable 3D Interaction-Trace World Model

30 篇
🔧 1 📖 1

MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

2 篇
🔧 10 📖 17

Learning to Assist: Collaborative VLAs for Implicit Human-Robot Collaboration

27 篇
⚡ 1 🔧 13 📖 14

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

28 篇
🔧 12 📖 15

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

27 篇
🔧 14 📖 16

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

30 篇
🔧 10 📖 18

Robots Need More than VLA and World Models

28 篇
🔧 2

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

2 篇
🔧 11 📖 13

VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training

24 篇
🔧 12 📖 15

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

27 篇
🔧 10 📖 19

World-Task Factorization for Robot Learning

29 篇
🔧 18 📖 9

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation

27 篇
— 2026 年 5 月 —
⚡ 1 🔧 8 📖 20

GEM: Generative Supervision Helps Embodied Intelligence

29 篇
🔧 13 📖 13

GEM: Generative Supervision Helps Embodied Intelligence

26 篇
🔧 9 📖 17

PhyPush: One Push is All You Need for Sensorless Physical Property Estimation with Physics-Guided Transformers

26 篇
⚡ 1 🔧 12 📖 17

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

30 篇
🔧 10 📖 11

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models

21 篇
📖 1

Bioinspired ionic thermoreceptors with anisotropic architecture for thermotactile perception in robots

1 篇
🔧 12 📖 11

Learning Structural Latent Points for Efficient Visual Representations in Robotic Manipulation

23 篇
⚡ 1 🔧 15 📖 26

Learning Structural Latent Points for Efficient Visual Representations in Robotic Manipulation

42 篇
🔧 9 📖 25

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR

34 篇
⚡ 1 🔧 9 📖 18

Key-Gram: Extensible World Knowledge for Embodied Manipulation

28 篇
🔧 11 📖 20

PhysBrain 1.0 Technical Report

31 篇
🔧 10 📖 23

SECOND-Grasp: Semantic Contact-guided Dexterous Grasping

33 篇
🔧 11 📖 17

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

28 篇
🔧 12 📖 15

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

27 篇
⚡ 1 🔧 14 📖 15

BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation

30 篇
🔧 10 📖 17

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts

27 篇
⚡ 1 🔧 7 📖 11

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

19 篇