OmniPiano: Diverse Dexterous Piano-Playing Challenges for Standard, Robust, Safe, and Multi‑Agent RL

Fan Xu1,*,†, Fan Yang2,*, Yifan Xu3,*, Ruiyuan Zhang4,*, Yimou Wu5, Nuo Yi6, Muning Wen2, Guang Chen7, Costas Spanos, Shangding Gu1,✉

1 University of California, Berkeley 2 Shanghai Jiao Tong University 3 The Chinese University of Hong Kong, Shenzhen 4 Japan Advanced Institute of Science and Technology 5 The Chinese University of Hong Kong 6 Nanjing University 7 Tongji University

* Equal contribution   † Project Lead   ✉ Corresponding author

OmniPiano overview: four tracks — Standard, Robust, Safe, and Multi-Agent RL — built around a shared piano-playing core
Figure 1: Overview of OmniPiano. A unified benchmark for learning dexterous piano playing.
1–5Shadow Hands
111max action dims
150songs
912task settings
36baseline algorithms
4LLM-based agents
StandardTwo Shadow Hands playing Für Elise.

Abstract

Dexterous robotic manipulation remains a major challenge in robotics, exemplified by piano playing, which requires coordinated control of fingers and provides a demanding testbed for reinforcement learning (RL). Although previous work established a benchmark for robotic piano playing, it is restricted to two hands and does not support unified evaluation. To address this gap, we introduce OmniPiano, a benchmark that provides one to up to five Shadow Hands and supports standard, robust, safe, and multi-agent RL within a shared piano-playing task family. OmniPiano is developed for scalable and diverse piano-playing tasks with configurable perturbations, explicit safety constraints, and decentralized cooperation settings. Particularly, to facilitate usability and extendability, OmniPiano adopts a highly modular design and provides comprehensive tasks to support RL study of task performance, robustness, safety, and cooperation in dexterous control. With at least 912 task settings and 36 baseline algorithms, OmniPiano further incorporates LLM-based agents to broaden the evaluation scope and reveal new insights. Extensive evaluations show that state-of-the-art RL algorithms and frontier LLM-based agents still struggle on challenging piano-playing tasks, even under standard RL settings, highlighting substantial room for improvement in dexterous control. The code, dataset, and tutorial are available at https://omnipiano.site.

Why piano playing?

A precise, measurable testbed for dexterous control

A MIDI score specifies exactly which keys to press, hold, and release, and when. Comparing intended and actual key activations gives a concrete, note-level measure of how well an RL policy controls a high-dimensional, contact-rich body.

Two simulated Shadow Hands playing a piano keyboard
  1. High-dimensional continuous control

    Anthropomorphic Shadow Hands with position control on an 88-key piano: 44 DoF for two hands, scaling to 111 action dimensions with five hands.

  2. Spatial & temporal coordination

    Fingers on every hand must hit different notes at the right time, without colliding with each other in a shared workspace.

  3. Diverse task settings

    150 songs × 8 hand settings, with restricted or unrestricted hand mobility, plus perturbation, safety, and multi-agent variants of the same task.

Related benchmarks

How OmniPiano compares

OmniPiano keeps the high-dimensional dexterous piano playing of RoboPianist and adds robust, safe, and multi-agent RL tracks on a single unified task setting, together with LLM-agent evaluation.

Comparison of OmniPiano with related reinforcement-learning benchmarks.
FeatureRoboPianistRobust-Gym.Safety-Gym.OmniPiano
Action dimension451–302–1723–111
Dexterous piano playing✓✗✗✓
Robust RL track✗✓✗✓
Safe RL track✗✓✓✓
Multi-agent RL track✗✓✓✓
Unified task setting✗✗✗✓
LLM-agent evaluation✗✗✗✓

One core task, four tracks

Vary one factor at a time

Rather than combining every factor in a single environment, OmniPiano keeps the core piano-playing task fixed and varies one aspect per track, so performance differences can be attributed to a single design choice. All tracks are exposed through consistent Gymnasium (single-agent) and PettingZoo (multi-agent) interfaces in MuJoCo, and work with Stable-Baselines3, RLlib, OmniSafe, and TorchRL.

01

Standard RL

How well can a policy play, as repertoire and the number of hands grow?

Hand settings. Eight hand settings from one to five Shadow Hands. Unrestricted hands move freely across the keyboard; restricted hands are each confined to a fixed region, which keeps exploration and credit assignment tractable as the number of hands grows.

Reward. Following RoboPianist, rt = rtkey + rtmatch − λenergy rtenergy, rewarding correct key presses, fingertip-to-key proximity, and low actuation. Two variants: a fingering-annotated reward, and an optimal-transport reward that needs no fingering annotations and so supports arbitrary MIDI files and hand configurations.

Baselines. PPO, SAC, CrossQ, TD3, TQC, and four LLM-based agents that optimize keyframes over ten iterations.

Results

Two-hand evaluation

RL & LLM leaderboard

Click a metric to sort

Include
The all-songs view uses a macro average across Twinkle Twinkle, Pictures at an Exhibition: Great Kiev, and Winter Wind; single-song views show that song directly. Tokens include prompt and completion tokens.
Rank Method
1CrossQRL1,172.3—
2TQCRL1,127.7—
3SACRL1,136.5—
4GPT-6 AstraLLM1,009.1582.3K
5PPORL1,121.2—
6Claude Opus 5LLM899.1362.8K
7Gemini 3.8 FlashLLM794.3416.8K
8DeepSeek V4.1 FlashLLM831.8938.2K
9TD3RL756.2—

Evaluation budget. LLM agents are evaluated over 10 keyframe-optimization attempts, while RL methods are trained for 5M environment steps. Token counts apply only to LLM methods. GPT-6 Astra and Claude Opus 5 each lack one recorded Twinkle attempt; their token averages use the nine available attempts for that song.

Bar charts of average evaluation reward and mean token usage per optimization attempt for LLM agents and RL methods on three piano pieces
Figure 5: Average evaluation reward (↑: top row, higher is better) and mean token usage (↓: bottom row, lower is better) per optimization attempt for LLMs across ten keyframe-optimization iterations and RL methods over 5M training steps in the two-hand setting on three piano pieces.
Radar charts of F1 score for PPO, SAC, DeepSeek-V4.1-Flash and GPT-6 Astra across eight hand settings on three songs
Figure 6: F1-score comparison of PPO, SAC, DeepSeek-V4.1-Flash, and GPT-6 Astra across eight hand settings (U = unrestricted, R = restricted) on three piano pieces.

Takeaway. More hands do not consistently improve F1. RL methods improve more stably, while LLM-based agents can achieve rapid F1 gains within a few attempts at substantial token cost; GPT-6 Astra shows the strongest iterative in-context learning. Restricted hand assignments converge faster and reach higher final returns than unrestricted ones.

More results: restricted vs. unrestricted hands
Evaluation reward for PPO and SAC under restricted and unrestricted configurations
Figure 10: Periodic evaluation reward during training for PPO and SAC under restricted and unrestricted configurations across the three songs.
02

Robust RL

Does the playing hold up when signals and physics are perturbed?

Robust RL design: action, environment, reward and observation disruptors acting on the agent–environment loop
Figure 2: Robustness perturbations are introduced through four independent disruptors acting on the agent–environment interaction loop.

Signal perturbations

Operators Do, Da, Dr corrupt observations, actions (actuation uncertainty), and rewards during interaction.

Environment perturbations

De changes gravity, fingertip–key contact friction, and the initial pose of each hand, per episode or per step.

Composable

Gaussian, uniform, or shift noise at low / medium / high levels on any channel, combined freely to test isolated and compound effects.

Results

PPO evaluation reward under six perturbation channels and three noise families
Figure 7: Periodic evaluation reward of PPO across six robustness channels under Gaussian, uniform, and shift perturbations, with the clean environment shown as a reference.

Takeaway. Action and hand-pose perturbations cause the largest performance drops, while contact friction and gravity are generally less disruptive. Across algorithms, A2PSAC and SAC stay higher and more stable, whereas EPPO is weakest overall.

More results: robust RL algorithms and composed perturbations
Cross-algorithm robustness under Gaussian action perturbations on Clair de Lune
Figure 12a: Cross-algorithm robustness (PPO, SAC, A2PSAC, SCPO, EPPO) on Clair de Lune under Gaussian action perturbations.
Cross-algorithm robustness under Gaussian reward perturbations on Clair de Lune
Figure 12b: The same comparison under Gaussian reward perturbations.
Composed Gaussian perturbations on Für Elise with three hands
Figure 13: Progressively composing observation, reward, action, gravity, and hand-pose perturbations on three-hand Für Elise.
03

Safe RL

Accurate playing is not the same as safe playing.

Safe RL design: four safety semantics combined with three cost settings
Figure 3: Four safety semantics, Joint Range, Actuator Power, Injured Finger, and Hand Collision, define the constrained behaviors. Event, Fraction, and Excess provide alternative measures of constraint violations.

Safety semantics

  • Joint Range: keep selected joints within prescribed intervals
  • Actuator Power: limit instantaneous mechanical power
  • Injured Finger: extra limits on a designated finger
  • Hand Collision: limit contact forces between hands

Cost settings

  • Event: does a violation occur at this step?
  • Fraction: share of monitored elements over their limits
  • Excess: normalized deviation beyond the threshold

Reward and cost stay separate: policies maximize musical return subject to an episodic cost limit, and evaluation assesses playing quality and constraint satisfaction jointly.

Results

Training reward and safety cost of CUP with two, three, and four hands, compared with unconstrained PPO
Figure 8a: Hand-count sensitivity of CUP: training reward and safety cost for 2–4 hands. Solid lines denote CUP; dashed lines denote unconstrained PPO. Three seeds (mean ± std); the dotted line marks the cost budget.

Takeaway. Increasing the number of hands can make safe RL more difficult: in this task, four-hand policies struggle more to reduce safety costs and achieve lower rewards than two- or three-hand policies.

More results: hand-count sensitivity of PPO-Lag
Training reward and safety cost of PPO-Lag with two, three, and four hands, compared with unconstrained PPO
Figure 8b: Hand-count sensitivity of PPO-Lag, with the same setup; dashed lines denote unconstrained PPO.
04

Multi-Agent RL

Split the hands across decentralized agents that share one musical goal.

Multi-agent RL design: base configuration and four SCHO variants
Figure 4: A typical four-hand example. Users can freely configure the MIDI score, number of hands and agents, hand-to-agent assignments, and agent-specific observation and action ranges and overlaps.

Scalability

How many agents control the same set of hands.

Coupling

How much agents' keyboard action ranges overlap.

Heterogeneity

Balanced vs. imbalanced hand assignments.

Observability

How much of the keyboard and teammates each agent sees.

Results

Team return of IPPO, MAPPO, HAPPO, A2PO and FACMAC across five cooperation settings
Figure 9: Task-wise training performance of IPPO, MAPPO, HAPPO, A2PO, and FACMAC on four-hand Winter Wind across the Base task and four SCHO variants, with centralized PPO as a reference. Compared with Base, Scalability and Observability generally improve performance, while Heterogeneity and especially Coupling make policy learning more difficult; Coupling most clearly differentiates the algorithms.

Takeaway. More agents can be beneficial: with the same four hands, performance improves from centralized PPO with one agent, to the two-agent Base setting, and further to the four-agent Scalability setting.

More results: team return and musical F1 per task
Detailed MARL results: team return and musical F1 per algorithm and task
Figure 23: Detailed MARL results. The first row compares task-wise team return per algorithm; the second and third rows compare team return and musical F1 across algorithms for each task.

Cross-benchmark insights

Five findings

1

Song duration alone does not explain difficulty

Note density (NPS) is more informative than episode length, while difficulty also depends on fingering complexity, hand coordination, and spatial key transitions.

2

More hands do not guarantee better performance

Extra hands do not consistently raise F1, and in safe RL four-hand policies struggle more to reduce costs: additional control freedom increases exploration and coordination difficulty.

3

RL improves steadily; LLM agents gain fast at a token cost

LLM-based agents can achieve rapid F1 gains within a few attempts, but with higher variance and substantial token use. GPT-6 Astra shows the strongest iterative in-context learning.

4

Action and hand-pose perturbations are most disruptive

They cause the largest drops, highlighting actuation and initialization as the key robustness challenges in dexterous piano playing.

5

More agents can help

With the same four hands, performance improves from centralized PPO, to two agents, to four agents with one hand each.

Taken together

State-of-the-art RL algorithms and frontier LLM-based agents still struggle on challenging piano-playing tasks, even under standard RL settings, leaving substantial room for improvement in dexterous control.

Rollouts

Videos

StandardOne hand, unrestrictedFür Elise
StandardTwo hands, unrestrictedFür Elise
StandardThree hands, restrictedPictures at an Exhibition (Great Kiev)
StandardFour hands, restrictedWinter Wind
Five Shadow Hands playing piano in the restricted setting
StandardFive hands, restrictedEach hand confined to its own keyboard region
RobustTwo hands, action noiseClair de Lune
RobustThree hands, observation + action noiseFür Elise
RobustFive hands, observation noiseFür Elise
Hand-collision safety task rollout
SafeHand collisionLimiting contact between neighboring hands
Multi-AgentTwo agents, four handsWinter Wind · Base cooperation setting
Multi-AgentOne hand vs. three handsWinter Wind · Heterogeneity setting
Multi-AgentFour agents, one hand eachWinter Wind · Scalability setting

Data

MIDI Dataset

OmniPiano uses the Piano Fingering (PIG) dataset. Due to licensing restrictions, PIG cannot be redistributed on GitHub, so download and preprocess it locally. The steps are identical to RoboPianist's.

  1. Download

    Register for a free account on the PIG website and download PianoFingeringDataset_v1.2.zip.

  2. Extract

    Unzip the archive and keep the PianoFingeringDataset_v1.2 folder; its path is used in the next step.

  3. Preprocess

    Convert the dataset into the format OmniPiano loads. This creates a pig_single_finger directory.

    robopianist preprocess --dataset-dir /PATH/TO/PianoFingeringDataset_v1.2
  4. Verify

    Check that preprocessing succeeded. It should print “PIG dataset is ready to use!”

    robopianist --check-pig-exists