Skip to content

VLA 線 · Vision-Language-Action

VLA 研究日報VLA 研究日报

VISION-LANGUAGE-ACTION · cs.RO + cs.AI + cs.LG


Vision-Language-Action(VLA)機器人系統 — 整合 cs.RO、cs.AI、cs.LG 三條 arxiv 流。 重點追蹤 flow matching、世界模型、具身推理等前沿方向,每日 09:00 CST 由 Qwen3.5-Plus 自動評級Vision-Language-Action(VLA)机器人系统 — 整合 cs.RO、cs.AI、cs.LG 三条 arxiv 流。 重点追踪 flow matching、世界模型、具身推理等前沿方向,每日 09:00 CST 由 Qwen3.5-Plus 自动评级

— 2026 年 9 月 —
今天
🔧 5 📖 31

Physics filtering favors the generalization of robot learning

36 篇
昨天
📖 9

Physics filtering favors the generalization of robot learning

9 篇
2 天前
🔧 5 📖 30

Physics filtering favors the generalization of robot learning

35 篇
3 天前
📖 10

Physics filtering favors the generalization of robot learning

10 篇
4 天前
📖 10

Physics filtering favors the generalization of robot learning

10 篇
5 天前
🔧 3 📖 37

Physics filtering favors the generalization of robot learning

40 篇
6 天前
🔧 6 📖 9

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

15 篇
7 天前
🔧 2 📖 7

RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning

9 篇
🔧 2 📖 8

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

10 篇
🔧 2 📖 13

IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

15 篇
— 2026 年 8 月 —
📖 5

ZEST: Zero-shot embodied skill transfer for athletic robot control

5 篇
🔧 5 📖 25

ZEST: Zero-shot embodied skill transfer for athletic robot control

30 篇
🔧 5 📖 23

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

28 篇
🔧 5 📖 33

ZEST: Zero-shot embodied skill transfer for athletic robot control

38 篇
🔧 5 📖 25

ZEST: Zero-shot embodied skill transfer for athletic robot control

30 篇
🔧 5 📖 15

Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model

20 篇
🔧 1 📖 2

Aerial tactile perching via an anthropomorphic hand with embodied soft tactile receptors

3 篇
🔧 9 📖 22

SONIC: Supersizing motion tracking for natural humanoid whole-body control

31 篇
⚡ 1 🔧 10 📖 24

SONIC: Supersizing motion tracking for natural humanoid whole-body control

35 篇
⚡ 1 🔧 24 📖 13

SONIC: Supersizing motion tracking for natural humanoid whole-body control

38 篇
⚡ 1 🔧 19 📖 11

SONIC: Supersizing motion tracking for natural humanoid whole-body control

31 篇
🔧 8 📖 11

SONIC: Supersizing motion tracking for natural humanoid whole-body control

19 篇
📖 2

SONIC: Supersizing motion tracking for natural humanoid whole-body control

2 篇
🔧 1 📖 2

Learning contact representations in real-world clutter for universal robotic grasping

3 篇
⚡ 1 🔧 10 📖 21

RoboSynChallenge: Mastering Real-World Dexterity via Generalizing Synthesized Manipulation Skills

32 篇
⚡ 1 🔧 7 📖 21

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

29 篇
🔧 7 📖 37

Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations

44 篇
🔧 18 📖 8

SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning

26 篇
⚡ 1 🔧 10 📖 19

Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection

30 篇
📖 1

Energy-Guided Flow Matching

1 篇
⚡ 1 🔧 19 📖 18

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

38 篇
🔧 15 📖 9

Kitchen Robotic Manipulation utilizing Foundation Models

24 篇
🔧 11 📖 14

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies

25 篇
🔧 19 📖 8

Towards General Language-Conditioned Latent Safety Filters

27 篇
🔧 5 📖 14

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

19 篇
🔧 1 📖 2

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

3 篇
⚡ 1 🔧 8 📖 28

It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation

37 篇
— 2026 年 7 月 —
🔧 10 📖 14

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

24 篇
🔧 9 📖 16

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

25 篇
⚡ 1 🔧 9 📖 16

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

26 篇
⚡ 1 🔧 8 📖 16

Agile perceptive multiskill locomotion for quadrupedal robots in the wild

25 篇
📖 1

Offline RL with Hierarchical Action Chunking

1 篇
🔧 6 📖 6

PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics

12 篇
⚡ 1 🔧 12 📖 11

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

24 篇
🔧 13 📖 11

STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models

24 篇
🔧 10 📖 16

ConceptTree: Bringing Semantic Transparency to Black-Box Decision Making for Robotic Manipulation

26 篇
⚡ 1 🔧 4 📖 17

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

22 篇
⚡ 1 🔧 9 📖 17

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

27 篇
🔧 15 📖 17

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

32 篇
🔧 9 📖 17

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

26 篇
🔧 10 📖 19

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

29 篇
🔧 10 📖 15

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

25 篇
⚡ 1 🔧 11 📖 18

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

30 篇
⚡ 1 🔧 11 📖 18

GeoProp: Grounding Robot State in Vision for Generalist Manipulation

30 篇
🔧 9 📖 23

Learning 4D Geometric Priors for Inference-Efficient World Action Models

32 篇
🔧 8 📖 20

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

28 篇
📖 1

High-resolution real-time mechanochromic tactile sensors

1 篇
📖 1

High-resolution real-time mechanochromic tactile sensors

1 篇