中文 English
🤖 LateAI 项目库 CV Zone 🦾 做自己的机器人 🦾 Build Your Own Robot
首页做自己的机器人 › ALOHA 双臂系统Build Your Own Robot/div › ALOHA Bimanual System (Detailed Guide)o">
🤲
🤲h1>ALOHA 双臂遥操作系统:硬件、标定与 ACTALOHA Bimanual Teleoperation System: Hardware, Calibration and ACT Traininghero-en">ALOHA — Bimanual TeleopeALOHA — Bimanual Teleoperation Hardware, Calibration & ACT Training"meta"> 📶 进阶 💰 About $20,000 for the original full systemDIY · 站内原创 🦾 DIY · Original on this site

📌 本分册目标

ALOHA(A Low-cost Open-soALOHA (A Low-cost Open-source Hardware System for Bimanual Teleoperation) is the bimanual teleoperation system Stanford released at RSS 2023. Its conclusion is blunt: 学会「about $20,000 of hardware子、抛接, and you can capture demonstration data good enough to learn fine two-handed actions like “threading a zip tie, inserting a battery, opening a cup, and tossing a ping-pong ball” — tasks that used to be considered possible only with expensive force-feedback teleoperation rigs.>论文里 4 个任务的成功率分别是 96% In the paper, the success rates of the 4 tasks are 本分册按 架构 → 预算 → 装配标定 →This guide walks through five stages: 训练architecture → budget → assembly & calibration → capture → training & evaluation也先讲清. Let's also be clear up front about how it differs from Project 04 (Koch): Koch teaches you how to “get started”, ALOHA teaches you how to “scale up” — dual arms, four cameras, 50Hz capture, every detail in service of fine manipulation.e">💡 本站分册为原创中文技术整理,基💡 This guide is an 开资料original Chinese-language technical write-up库图文与, written from the paper and official public materials, and does not reproduce the original repository's images or code. ALOHA was open-sourced by Tony Zhao et al. at Stanford, and the commercial kit is provided by Trossen Robotics; Interbotix / ViperX / WidowX are their product line names, and RealSense is an Intel trademark. This site is not affiliated with any of the above.section class="sec">

🏗️ 第一段:系统架构(4 臂 + 4 相机)<🏗️ Part 1: System architecture (4 arms + 4 cameras)">ALOHA 的「主从同构」与 Koch 是同一个ALOHA's “leader-follower isomorphism” is the same idea as Koch's, but scaled up to 主臂同dual arms + a larger reach步复现整. You hold one leader arm in each hand and drag them in sync, and the two follower arms reproduce the entire two-handed cooperative motion.table"> 部件角色 Role WidowX 250 S ×2WidowX 250 S ×2度,机身较轻、便于Leader arm. 6 degrees of freedom; the body is lighter and easier to drag by hand.td>ViperX 300 S ×2ViperX 300 S ×2自由度,臂展 75Follower arm. 6 degrees of freedom, 750mm reach, up to about 1500mm between the two arms; both rigidity and workspace are larger.td>RealSense D405 ×4控制主机官方可选预配置笔记本(Control hostem76 ServThe official option is a preconfigured laptop (System76 Serval WS: i9 + 32GB + RTX 4060) or a mini PC; for a self-built machine, a discrete GPU + Ubuntu 22.04 is recommended.>

腕部相机是这套系统能做精细任务的关键。固The wrist cameras are the key to this system's ability to do fine tasks.腕部视角Fixed camera positions can't see the contact details between gripper and object; only the wrist view provides the decisive information about “the relative relationship between fingers and workpiece”. The official project also provides 3D-printed mounts for the top / bottom / wrist cameras — don't use a makeshift clamp-on mount, because view stability directly affects training results.n">⚠️ 官方文档中 Mobile 机型的相机配置在⚠️ In the official documentation the Mobile variant's camera configuration is described inconsistently between the spec sheet and the packing list (one place says 3× D405, the packing section says 3× USB Camera). Confirm with the vendor before ordering. A mistake in this kind of detail will stall the capture stage outright.section class="sec">

💸 第二段:预算三条路线怎么选

💸 Part 2: How to choose among the three budget routesin">
  • 整机路线(Stationary / MoFull-system route):约 (Stationary / Mobile): about the $20,000 level, including 4 arms + 4 cameras + host + fixtures. Suited to teams with ample funding that need data immediately.路线(单臂对):1 主 + 1Solo route相机。用 (a single arm pair): 1 leader + 1 follower + 2 cameras. Use it to get the whole “capture → training → evaluation” pipeline working end to end; the cost drops sharply, and it also verifies whether your team really needs dual arms.线:先走 LeRobot 生态Low-cost DIY route臂(可参: start with a low-cost arm from the LeRobot ecosystem (see Project 04 Koch on this site), get the algorithms and data pipeline down pat, and only move up to ALOHA-class hardware once the task really needs bimanual cooperation.lass="lead">这三条路的共同点是算法侧完全一样—What the three routes share is that 有几条the algorithm side is exactly the same决策顺序 — ACT doesn't care how many arms you have. So the real decision order should be: first confirm whether the task needs bimanual cooperation, then decide how much to spend.ction class="sec">

    🔧 第三段:装配与标定

    1
    先装工装与相机支架,再上机械臂
    顺序反了会被线缆绊住。顶视机位要能完整覆盖双臂工Reverse the order and you'll trip over cables. The top view must fully cover the dual-arm workspace, and the bottom view must see under the grippers; the wrist mount must be
    2
    按官方尺寸布置臂间距与走线
    从臂间距按官方尺寸(双臂最远约 1500mm)布Space the follower arms according to the official dimensions (up to about 1500mm apart). Route cables to fixed points and leave slack for movement — being pulled taut during motion directly causes position error.div class="step">
    3
    按官方 Bringup 流程上电自检
    Power up and self-check following the official Bringup procedure">确认电源(12V 20A 与 12V 10A)与Confirm that the power supplies (12V 20A and 12V 10A) and the USB hub are all recognized, then test-turn each joint and check direction, limits, and binding. Wrist joints are the most likely to loosen in transit.div class="step">
    4
    主从标定:同姿态记录偏置
    把主臂与从臂摆到同一姿态,记录各关节偏置作为标定Put the leader and follower arms into the same pose and record each joint's offset as the calibration baseline. The typical symptom of not calibrating is the follower being “off by one angle”, after which the dataset is basically worthless.div class="step">
    5
    重力补偿不能省
    Stationary 机型有硬件重力补偿器(弹簧The Stationary variant has hardware gravity compensators (springs / counterweights) and the Mobile variant uses software gravity compensation. Without it, the follower arms sag under their own weight, and b>。the captured data is full of useless “fighting gravity” actions .div class="step">
    6
    遥操作预演:先把手感调顺
    双手握主臂慢慢动,从臂应几乎同步跟随;有延迟或抖Hold the leader arms in both hands and move them slowly; the followers should follow almost in sync. If there's lag or jitter, check USB bandwidth and bus load first. Try gripper opening/closing and wrist rotation too — when you're capturing data there's no time to adjust hardware.ection>

    🎥 第四段:50Hz 数据采集流程

    🎥 Part 4: The 50Hz data capture pipeline">官方流程可分为四步,建议照这个顺序走,不要跳步:The official pipeline has four steps; follow this order and don't skip steps:n">
  • Task Creation:定义任务与物Task Creation机化位置: define the task and the object placement rules (for example, randomizing position along a reference line).ding Episodes:用Recording Episodes双手始终: use a 主臂,foot pedal贯。:回放检Episode Playback时间对齐: play back and check, confirming that images, joint states, and actions are time-aligned.Recording:自动化批量Auto-Recording> : automated batch capture for higher volume.le class="ptable"> 参数取值与理由 Value and rationale 采集频率50Hz。足够捕捉精细Capture rate奏,又不至于让数据50Hz. Enough to capture the rhythm of fine motions without letting the data volume get out of control.td>chunk size90(ACchunk size一次预测一整段,是90 (ACT's action-chunk length). Predicting a whole chunk at once is the source of policy stability.td>单条长度约 600~1000 Episode length一次完整的双手操作About 600–1000 steps, corresponding to one complete two-handed operation.td>演示条数官方实践里每个任务约 Number of demonstrations(约 10 分钟数In official practice about 50 per task (roughly 10 minutes of data) is enough to produce a usable policy; fine tasks can use a few more.>
    💡 光照要稳定、物体位置要随机化、失败演示要单独💡 Keep lighting stable, randomize object positions, and label failed demonstrations separately — these three are the most common “I only found out afterwards that it was wrong” issues at the capture stage. Failed samples are used at the evaluation stage to judge the policy's ability to recover, so don't delete them all.section class="sec">

    🧠 第五段:ACT 为什么能扛住精细操作模仿学习最经典的问题是误差累积:逐The classic problem of imitation learning is 点,下error accumulation的状态就: when predicting actions step by step, each step drifts a little, and the next step's input state is already off, so it wanders further and further. ACT (Action Chunking with Transformers) solves this with 块,让action chunking长的时间 — predicting an entire future chunk of actions at once, letting the policy make decisions on a longer time scale, which greatly weakens the accumulation effect.n">
  • 结构:以条件 VAE 的形式训练;编码器Architecture作序列压: trained as a conditional VAE; the encoder compresses the action sequence into a style variable, and the decoder fuses multi-view images and joint positions to output the whole chunk of actions.>:部署时丢掉编码器、把 style Inference z 置: at deployment time, drop the encoder and set the style variable z to zero (take the prior mean) — this detail determines real-time performance, and it's also the source of the reproduction gap many people see.>:在独显主机上跑;先在小数据集上确认Training拟合(l: run it on a machine with a discrete GPU; first confirm on a small dataset that it can overfit (loss clearly drops), then scale up to the full dataset.>:先用固定初始位置,再逐步随机化;每Evaluation务至少 : start with fixed initial positions, then gradually randomize; run at least 20 trials per task — don't draw conclusions from single-digit success rates.lass="lead">排查顺序的建议:表现差时先怀疑视觉(机位Suggested troubleshooting order标定,其: when performance is poor, suspect vision (camera placement / lighting) and calibration first, and the network architecture only after that. In practice, the vast majority of “can't train it” cases come from the first two.ction class="sec">

    🚧 避坑要点

    • 腕部相机别省:它是精细任务的胜负手,先保Don't skimp on the wrist cameras定再谈其: they are the deciding factor in fine tasks; get the wrist views stable before anything else.须做:差一个角度,几十条演示就Leader-follower calibration is mandatory。:否则从臂下Gravity compensation (hardware or software) can't be skippedli> : otherwise the follower arms sag and the demonstrations are full of compensating actions.定:窗帘、顶灯的变化会被策略当Keep ambient light stable,换个时: changes in curtains and ceiling lights get learned as features by the policy, and stop working at a different time of day.o 或低配再上整机:算法流程与Start with Solo or a low-cost build before going full-system流程跑熟: the algorithm pipeline is independent of hardware tier, so learning the pipeline on cheap hardware first is the most economical route.试验评估:精细操作任务方差大,Don't evaluate with single-digit trials计意义。: fine manipulation tasks have high variance; 20+ runs are needed for statistical meaning./b>:代码与硬件资料按官方许可开放,Copyright reminderTros: code and hardware materials are open under the official license, and the kit is a Trossen product; when citing the paper and figures follow academic norms, and read the respective license terms before commercial use.ion>

      🔗 官方资料直达

      • ALOHA 官方项目主页GitHub · tonyzhaozh/alohGitHub · tonyzhaozh/aloha遥操作、数据采集、回放与自动采集脚本Trossen 官方 ALOHA 文档ViperX-300 规格页📄 ACT 论文(RSS 2023)💻 ACT 代码库
      • The official implementation of the action-chunking algorithma href="https://mobile-aloha.github.io/"" target="_blank" rel="noopener">🚙 Mobile ALOHALeRobot × ALOHA 官方指南LeRobot × ALOHA official guide开源生态下的低成本复现路径

        ❓ 常见问题

        ALOHA 和 Koch 到底该选哪个?

        按任务选:单臂抓取 / 分拣、预算几千元

        能不能只买一套主从臂(Solo)?

        可以,官方就提供 Solo 配置(1 主 1 从Yes. The official project offers a Solo configuration (1 leader + 1 follower + 2 cameras). It's good for validating the pipeline and doing single-arm tasks, but it can't do two-handed cooperation — “dual arms” is the reason ALOHA exists, so paying this budget for single-arm tasks alone is a waste.>

        为什么一定要 4 个相机?少几个行不行?

        最少要保证「全局视角 + 腕部近景」两个信息层级At minimum you must have two information levels: “global view + wrist close-up”. Remove the wrist cameras and the failure rate of fine contact tasks rises noticeably, because that amounts to asking the policy to control contact “blindfolded”. Fixed camera positions plus wrist cameras is the configuration validated on the paper's tasks.>

        硬件这么贵,值得吗? The hardware is expensive — is it worth it?ass="faq-a">

        先问自己一个问题:你的任务是否真的需要双臂Ask yourself one question first: does your task 用便宜really need bimanual cooperation达到同样. If not, a solution an order of magnitude cheaper can achieve the same result; if it does (threading zip ties, two-handed cooperative assembly, etc.), then ALOHA is currently the most complete and best-validated public reference, and it can save you a lot of trial and error.>

        采集要多久、训一次要多久? How long does capture take, and how long does one training run take?ass="faq-a">

        官方实践里,每个任务约 50 条演示(约 10 In official practice, about 50 demonstrations per task (roughly 10 minutes of data) is enough. Data capture can usually be finished within a day; training takes a few hours on a machine with a discrete GPU, but 问题、iteration据,往往 is the truly time-consuming part — evaluation, finding problems, and collecting more data often take several rounds.>

        ✅ 下一步

        回到「做自己的机器人」栏目项目 06,那里有本分Go back to Project 05 in the “Build Your Own Robot” section, where you'll find this guide's entry point, the official code repository, the ACT algorithm library, and direct cards for Mobile ALOHA. Suggested order: read the paper and official docs to understand the pipeline → decide on a budget route (Solo / full system / low-cost first) → assemble and calibrate → capture 50 demonstrations → train ACT → evaluate on the real robot 20+ times.ions"> 📦 打开 ALOHA 代码库 📦 Open the ALOHA code repository btn-ghost" href="/index.html#diy"">← 返回「做自己的机器人」