Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation

Jun Hyuk Lee 1 Chihyeong Lee 1 Jooeun Ahn 1,2,3,*
1Department of Physical Education, Seoul National University, Seoul, Republic of Korea 2Institute of Sport Science, Seoul National University, Seoul, Republic of Korea 3SNU Robotics Institute, Seoul National University, Seoul, Republic of Korea
*Corresponding author: ahnjooeun@snu.ac.kr
slow (×1/4)
slow (×1/4)

Abstract

The massive overactuation in the human musculoskeletal system makes it challenging to train musculoskeletal models to generate human-like motion via reinforcement learning, primarily because exploration in the resulting high-dimensional and redundant action space is extremely inefficient. To address this problem, we propose the $\lambda$-hold controller, inspired by the equilibrium-point (EP) hypothesis, which has been widely supported by extensive evidence from human motor control studies. The policy's control variable is the per-muscle EP threshold length $\lambda$, from which a stretch-reflex recruitment law computes the muscle excitations automatically. Holding each $\lambda$ over an interval of the gait phase also sharply reduces the frequency at which the policy must be queried. Consequently, the controller, to our knowledge for the first time, enables a muscle-actuated skeletal model to learn human-like sprinting using only a minimal reward within an hour of training. The efficient exploration through the proposed $\lambda$-hold controller is not merely an engineering trick but an approach grounded in physiology, bringing together the EP hypothesis, intermittent control, and optimal feedback control. Beyond encapsulating human-like behavior in predictive simulation, this achievement contributes to developing a learnable model of the human motor controller.

Overview

Overview of the lambda-hold controller

Ours vs. best baseline (DEP-RL)

λ-hold (Ours)
Sagittal
Frontal
DEP-RL (best baseline)
Sagittal
Frontal

Both controllers were trained for the same 150M simulation steps. Our λ-hold controller, however, reaches that same number of simulation steps in far less wall-clock time than DEP-RL.

Results & Comparisons

λ-hold (Ours)
Training progression
Plain SAC
Excitation-Hold
DEP-RL
Synergy
Start of training
75M
simulation steps
150M
simulation steps

BibTeX

@article{placeholder2026,
  title   = {Your Project Title Here},
  author  = {Author One and Author Two and Author Three},
  journal = {Venue / Conference},
  year    = {2026}
}