VLA 線 · Vision-Language-Action
VLA 研究日報VLA 研究日报
VISION-LANGUAGE-ACTION · cs.RO + cs.AI + cs.LG
Vision-Language-Action(VLA)機器人系統 — 整合 cs.RO、cs.AI、cs.LG 三條 arxiv 流。 重點追蹤 flow matching、世界模型、具身推理等前沿方向,每日 09:00 CST 由 Qwen3.5-Plus 自動評級Vision-Language-Action(VLA)机器人系统 — 整合 cs.RO、cs.AI、cs.LG 三条 arxiv 流。 重点追踪 flow matching、世界模型、具身推理等前沿方向,每日 09:00 CST 由 Qwen3.5-Plus 自动评级。
Physics filtering favors the generalization of robot learning
Physics filtering favors the generalization of robot learning
Physics filtering favors the generalization of robot learning
Physics filtering favors the generalization of robot learning
Physics filtering favors the generalization of robot learning
Physics filtering favors the generalization of robot learning
GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
ZEST: Zero-shot embodied skill transfer for athletic robot control
ZEST: Zero-shot embodied skill transfer for athletic robot control
$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
ZEST: Zero-shot embodied skill transfer for athletic robot control
ZEST: Zero-shot embodied skill transfer for athletic robot control
Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model
Aerial tactile perching via an anthropomorphic hand with embodied soft tactile receptors
SONIC: Supersizing motion tracking for natural humanoid whole-body control
SONIC: Supersizing motion tracking for natural humanoid whole-body control
SONIC: Supersizing motion tracking for natural humanoid whole-body control
SONIC: Supersizing motion tracking for natural humanoid whole-body control
SONIC: Supersizing motion tracking for natural humanoid whole-body control
SONIC: Supersizing motion tracking for natural humanoid whole-body control
Learning contact representations in real-world clutter for universal robotic grasping
RoboSynChallenge: Mastering Real-World Dexterity via Generalizing Synthesized Manipulation Skills
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations
SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning
Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection
Energy-Guided Flow Matching
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
Kitchen Robotic Manipulation utilizing Foundation Models
ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies
Towards General Language-Conditioned Latent Safety Filters
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation
CG-World: A Large-Scale World-State Dataset and Protocol for World Models
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
Agile perceptive multiskill locomotion for quadrupedal robots in the wild
Offline RL with Hierarchical Action Chunking
PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics
NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation
STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models
ConceptTree: Bringing Semantic Transparency to Black-Box Decision Making for Robotic Manipulation
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration
FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action
GeoProp: Grounding Robot State in Vision for Generalist Manipulation
Learning 4D Geometric Priors for Inference-Efficient World Action Models
Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning
High-resolution real-time mechanochromic tactile sensors
High-resolution real-time mechanochromic tactile sensors