Achieving human-like dexterity is the next frontier for robotics, and yet dexterity data is often subtly hard to scale. Real-world dexterity data, including things like finger-pose estimates, is often slightly off, making it physically invalid and hard to execute on real hardware and hard to learn from.
DO AS I DO is an algorithm for reconstructing and retargeting monocular RGB videos to robot hands, outperforming the state of the art and even working from generated videos. @bhawna_paliwal_@HarithejaE@willjhliang and @notmahi join us to tell us more.
Watch Episode 107 of RoboPapers, with @chris_j_paxton and @DJiafei, now!
Grounded API is live.
- SOTA on hand-tracking benchmarks (< 1 cm)
- SOTA on SLAM benchmarks
- In-the-wild ego data -> enriched data in minutes
- Integration with @huggingface@LeRobotHF & @rerundotio
- Built for @BitRobotNetwork RoboCap suite
Technical report & more↓
I didn't get what GPT-6 meant for robotics until I actually tried it.
GPT-6 Astra just does physical ICL out of the box.
we drop a recording of a human doing a novel task into 𝗰𝗼𝗱𝗲𝘅 app. Prompt it to drive a robot arm the same way.
It just works on the first pass!
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects.
Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
Excited to share SPD: simulation pre-training for dexterity. We pre-trained a policy in simulation and fine-tuned with less than 2 hours of real data (with @sarthakkamat)
Your policy doesn't need 7B params. It simply needs dense features.
Introducing Patch Policy: pretrained ViT + small transformer beats OpenVLA-OFT with 0.7% of its params, and trains on a 5090.
Here it inserts a cable (~2mm tol), and does it again as we unplug mid-rollout. 🧵
Pretrained ViTs see the world in rich, dense detail. Most policies pool it to a single vector before acting, discarding most of it.
We introduce Patch Policy: a minimal architectural extension that enables transformer-based policies to consume dense tokens directly, no billion-param VLM required. It outperforms a fine-tuned 7B VLA by 18% with ~0.7% of its parameters, enabling robust, precise manipulation.
I'm at #RSS2026 – presenting MolmoSpaces on Tuesday & CAP (Contact Anchored Policies) on Wednesday.
I've been thinking a lot about what robotics 2-5 years from now looks like: beyond teleop and position control. If you're interested about anything from robot free/human data or sim evals to force/torque controlled dexterous hands, let's chat! 🧵
How can generalist policies adapt to new challenges at deployment using skills they already have?
We optimize VLA *prompt inputs* with reinforcement learning, enabling efficient real-robot adaptation on complex tasks where existing methods struggle. 🧵
semantic-action-rl.github.io
🤖 How can we teach dexterous robots to perform precise, contact-rich assembly?
Introducing Play2Perfect: first learn to play with objects, then perfect the policy for tight insertion, multi-part assembly, and screwing.
Sound on! 🔊
🧵👇
Excited to share Do as I Do! We turn everyday human videos into physically consistent robot data that can be directly executed in the real world.
This was a fun collaboration with @bhawna_paliwal_ and @willjhliang, with lots of moving parts. More details in Mahi's thread below👇
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data?
Introducing Do as I Do: establishing a much needed correspondence between human videos and
Excited to release Do As I Do: a pipeline that turns everyday RGB human videos into dexterous robot manipulation trajectories!
Most prior work has been narrow, consisting of just lab recorded demos, egocentric-only, or assuming a closed set of objects. We develop a modular pipeline that can handle Internet, egocentric, exocentric, AND generated videos with virtually any rigid object. Also check out Mahi's post below!
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data?
Introducing Do as I Do: establishing a much needed correspondence between human videos and
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data?
Introducing Do as I Do: establishing a much needed correspondence between human videos and
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data?
Introducing Do as I Do: establishing a much needed correspondence between human videos and
Introducing Do as I Do 👀, a framework to transform everyday human videos into 100s of dexterous robot demos. Co-led with @bhawna_paliwal_ and @HarithejaE, and check out @notmahi's thread!
Here’s a little preview of our dexterous manipulation results. More about how we produce them from human reconstructions in this mini-thread! 🧵
x.com/notmahi/status…
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data?
Introducing Do as I Do: establishing a much needed correspondence between human videos and
160 Followers 84 FollowingEmbodied AI · World Models · Spatial Intelligence
Tracking papers, demos, and ideas shaping the next generation of intelligent machines.
77 Followers 672 FollowingGlobal Precision CNC Specialist | Master in Mechanical Engineering
Aerospace/Medical Custom Machining & 3D Printing | Prototype to Mass Production
6K Followers 199 FollowingSharpa is an AI robotics company dedicated to developing ultra-high performance robots and core components.
We Manufacture Time by Making Robots Useful.
1K Followers 902 FollowingIncoming asst professor at UIUC MechSE. My research focuses on learning dexterity from humans as well as building low-cost LEAP Hands.
500 Followers 1K FollowingThinking about low-prob estimation for AI safety, diffusion models, and ML for health. Postdoc at @Columbia. Prev: ML PhD at @NYU_Courant, @MSFTResearch.