wsagi commited on
Commit
9d06a74
Β·
verified Β·
1 Parent(s): b711ecd

docs: add 30s dance section + AMP-vs-PPO note

Browse files
Files changed (1) hide show
  1. README.md +14 -3
README.md CHANGED
@@ -19,7 +19,7 @@ pipeline_tag: reinforcement-learning
19
 
20
  # MimicKit-G1-LAFAN
21
 
22
- 4 motion-tracking policies for Unitree G1 (29-DoF) trained with [MimicKit](https://github.com/xbpeng/MimicKit) DeepMimic-style PPO on 15-second LAFAN1 retargeted slices: **fight, run, dance, jumps**. Single 4090 24G, ~1 h per motion, 4096 envs Γ— 1500 PPO iters.
23
 
24
  **Project repo (源码 / 倍现ε…₯口):**
25
  - πŸ§ͺ [vitorcen/isaaclab-experience](https://github.com/vitorcen/isaaclab-experience) β€” MimicKit training scripts, eval chain, USD material fix, design docs
@@ -34,13 +34,18 @@ pipeline_tag: reinforcement-learning
34
  | **dance** | 14.70 s / **98.0 %** | 🟒 触鑢 | LAFAN `dance1_subject1` [1746:2196] |
35
  | **jumps** | 14.70 s / **98.0 %** | 🟒 触鑢 | LAFAN `jumps1_subject1` [3441:3891] |
36
  | **run** | 9.45 s / **63.0 %** | 🟑 plateau | LAFAN `run1_subject2` [3341:3791] |
 
37
 
38
- 3/4 motions reach ship quality at iter 1500. `run` plateaus at ~63 %, likely needing ADD-style residual or curriculum sequencing β€” kept as a baseline.
39
 
40
  ---
41
 
42
  ## Demos (eval, 4 envs, with restored per-link material)
43
 
 
 
 
 
44
  <video controls width="640" src="https://huggingface.co/wsagi/MimicKit-G1-LAFAN/resolve/main/videos/fight.mp4"></video>
45
 
46
  <video controls width="640" src="https://huggingface.co/wsagi/MimicKit-G1-LAFAN/resolve/main/videos/dance.mp4"></video>
@@ -72,6 +77,7 @@ Slicing rule: pick the most representative central segment of each `*1` clip (sk
72
  MimicKit-G1-LAFAN/
73
  β”œβ”€β”€ README.md
74
  β”œβ”€β”€ videos/
 
75
  β”‚ β”œβ”€β”€ fight.mp4 (6.9 MB)
76
  β”‚ β”œβ”€β”€ run.mp4 (2.6 MB)
77
  β”‚ β”œβ”€β”€ dance.mp4 (1.9 MB)
@@ -83,6 +89,10 @@ MimicKit-G1-LAFAN/
83
  β”œβ”€β”€ run_15s/ …
84
  β”œβ”€β”€ dance_15s/ …
85
  β”œβ”€β”€ jumps_15s/ …
 
 
 
 
86
  └── assets/
87
  └── g1_textured.usd # 26 MB, per-link material restored
88
  ```
@@ -135,7 +145,8 @@ scripts/mimickit_eval_chain.sh
135
  ## Known limitations
136
 
137
  - **`run` plateau at 63 %**: the LAFAN run clip has fast contact + slip; vanilla DeepMimic reward + fixed action std saturates here. Likely fixes: ADD residual, motion-curriculum from walk, or larger action std at start.
138
- - **15 s only**: each policy is single-clip overfit; no multi-motion conditioning. For composition, see ProtoMotions / OmniH2O.
 
139
  - **No sim-to-real transfer attempted**: trained in Isaac Lab with raw observations, no domain randomization, no actuator delay model.
140
 
141
  ---
 
19
 
20
  # MimicKit-G1-LAFAN
21
 
22
+ 5 motion-tracking policies for Unitree G1 (29-DoF) trained with [MimicKit](https://github.com/xbpeng/MimicKit) DeepMimic-style PPO on LAFAN1 retargeted slices: **fight, run, dance, jumps** (15 s each), plus a longer **30 s dance** (`dance1_subject2`) that warm-starts from the 15 s dance and holds full-horizon tracking. Single 4090 24G, ~1 h per 15 s motion, 4096 envs.
23
 
24
  **Project repo (源码 / 倍现ε…₯口):**
25
  - πŸ§ͺ [vitorcen/isaaclab-experience](https://github.com/vitorcen/isaaclab-experience) β€” MimicKit training scripts, eval chain, USD material fix, design docs
 
34
  | **dance** | 14.70 s / **98.0 %** | 🟒 触鑢 | LAFAN `dance1_subject1` [1746:2196] |
35
  | **jumps** | 14.70 s / **98.0 %** | 🟒 触鑢 | LAFAN `jumps1_subject1` [3441:3891] |
36
  | **run** | 9.45 s / **63.0 %** | 🟑 plateau | LAFAN `run1_subject2` [3341:3791] |
37
+ | **dance (30 s)** | full 30 s, `Test_Return` **244** (>15 s base 227) | 🟒 longer-horizon | LAFAN `dance1_subject2` [1521:2421] |
38
 
39
+ 3/4 of the 15 s motions reach ship quality at iter 1500. `run` plateaus at ~63 %, likely needing ADD-style residual or curriculum sequencing β€” kept as a baseline. The **30 s dance** doubles the horizon (900 frames @ 30 fps): warm-started from the 15 s dance ckpt and run to 2500 iters, its converged `Test_Return` (244) actually exceeds the 15 s dance (227) β€” the discounted return saturates near the same ceiling regardless of clip length, so matching/exceeding it means full-horizon coverage held.
40
 
41
  ---
42
 
43
  ## Demos (eval, 4 envs, with restored per-link material)
44
 
45
+ **30 s dance (longer-horizon, `dance1_subject2`)** β€” DeepMimic PPO holds the full 30 s; an AMP baseline on the same clip could not keep rhythm (see notes below).
46
+
47
+ <video controls width="640" src="https://huggingface.co/wsagi/MimicKit-G1-LAFAN/resolve/main/videos/dance1s2.mp4"></video>
48
+
49
  <video controls width="640" src="https://huggingface.co/wsagi/MimicKit-G1-LAFAN/resolve/main/videos/fight.mp4"></video>
50
 
51
  <video controls width="640" src="https://huggingface.co/wsagi/MimicKit-G1-LAFAN/resolve/main/videos/dance.mp4"></video>
 
77
  MimicKit-G1-LAFAN/
78
  β”œβ”€β”€ README.md
79
  β”œβ”€β”€ videos/
80
+ β”‚ β”œβ”€β”€ dance1s2.mp4 (0.7 MB, 30 s dance, first 8 s)
81
  β”‚ β”œβ”€β”€ fight.mp4 (6.9 MB)
82
  β”‚ β”œβ”€β”€ run.mp4 (2.6 MB)
83
  β”‚ β”œβ”€β”€ dance.mp4 (1.9 MB)
 
89
  β”œβ”€β”€ run_15s/ …
90
  β”œβ”€β”€ dance_15s/ …
91
  β”œβ”€β”€ jumps_15s/ …
92
+ β”œβ”€β”€ dance_30s/ # longer-horizon dance (dance1_subject2)
93
+ β”‚ β”œβ”€β”€ model.pt # 11 MB, final (2500 iters, warm-started)
94
+ β”‚ β”œβ”€β”€ env.yaml # 900-frame / 30 s slice
95
+ β”‚ └── motion.pkl
96
  └── assets/
97
  └── g1_textured.usd # 26 MB, per-link material restored
98
  ```
 
145
  ## Known limitations
146
 
147
  - **`run` plateau at 63 %**: the LAFAN run clip has fast contact + slip; vanilla DeepMimic reward + fixed action std saturates here. Likely fixes: ADD residual, motion-curriculum from walk, or larger action std at start.
148
+ - **Single-clip overfit**: each policy tracks one clip; no multi-motion conditioning. For composition, see ProtoMotions / OmniH2O.
149
+ - **AMP vs phase-tracking on dance**: an AMP (adversarial motion prior) baseline on the full 131 s `dance1_subject2` failed β€” the discriminator plateaued at ~0.98 agent-accuracy and the policy could not keep the choreography's rhythm. Phase-conditioned DeepMimic tracking (used here) is the right tool for high-fidelity dance; AMP fits continuous/loopable skills that don't require exact timing. The 30 s dance is the DeepMimic answer to "longer dance."
150
  - **No sim-to-real transfer attempted**: trained in Isaac Lab with raw observations, no domain randomization, no actuator delay model.
151
 
152
  ---