PAKT: Physically-Aligned Kinesthetic Teaching for Reinforcement Learning

A kinesthetic teaching interface and high-frequency controller stack for contact-rich real-world reinforcement learning.

Abstract

Real-world reinforcement learning systems still struggle with contact-rich industrial manipulation, where tasks require micrometer-level precision, high success rates, human-level cycle times, and guidance that operators can provide safely and intuitively.

PAKT combines constrained kinesthetic teaching with a reference-generator and high-frequency impedance-control stack so human demonstrations remain feasible for the learned policy. Across four insertion and assembly benchmarks, it reduces cycle time by 23-48% and human intervention by 60-85%.

Video

Project video showing the PAKT workflow for contact-rich robot learning.

PAKT combines direct kinesthetic guidance with the same constrained execution stack used by the learned policy.

Contributions

Concrete system contributions for physically aligned kinesthetic teaching in real-world reinforcement learning.

Policy-feasible kinesthetic teaching

Operator-applied forces are routed through constrained admittance control, allowing direct physical guidance while shaping demonstrations with the same action parameterization and motion limits used during autonomous rollout.

High-frequency execution stack

Cartesian delta commands are issued at policy rate, shaped by a third-order reference generator, and tracked by a compensated 1 kHz Cartesian impedance controller.

Real-world ablation and evaluation

The evaluation separates controller-stack effects from kinesthetic-guidance effects across four contact-rich insertion and industrial assembly tasks.

Control Path

Switch between a sampled policy target and sinusoidal human-input wrench, then inspect how the selected signal passes through reference generation, impedance control, and the simulated end effector.

Policy signal

External wrench

Admittance Controller

Admittance controller

x¨ ad = Ma -1 ( fext - Da x˙ ad - Ka x˜ ad ) ead = Log ( Tad T EE -1 ) x˙ ad = ∫ x¨ ad dt , Tad = exp ( V^ Δt ) T ad prev , V^ = [ [ ω ] x v 0 0 ]

Paper notation: K_a is set to 0 for free Cartesian guiding.

Switch

Reference Generator

Reference generator

Tπ = TEE Ta , x˙ π = kv Log ( Tπ T EE -1 ) , x¨ π = 0 ( Tref,d , x˙ ref,d , x¨ ref,d ) = { ( Tad , x˙ ad , x¨ ad ) σ=1 ( Tπ , x˙ π , x¨ π ) σ=0 x⃛ ref = a2 ( x¨ ref,d - x¨ ref ) + a1 ( x˙ ref,d - x˙ ref ) + a0 Log ( Tref,d T ref -1 ) x⃛ ref = ω3 Log ( Tref,d T ref -1 ) + 3ω2 ( x˙ ref,d - x˙ ref ) + 3ω ( x¨ ref,d - x¨ ref ) x⃛ ref = clip ( x⃛ ref , x⃛ lim ) , x¨ ref = clip ( x¨ ref , x¨ lim ) , x˙ ref = clip ( x˙ ref , x˙ lim )

The paper chooses a triple pole at -omega for this critically damped filter.

Reference output

Impedance Controller

Impedance controller

τimp = J ( q ) T ( Λ x¨ ref + Dx ( Λ , Kx ) x˙ ˜ + Kx x˜ - Λ J˙ ( q˙ ) q˙ + 0.5 Λ˙ x˙ ˜ ) τnull = N ( Dnull q˙˜ + Knull q˜ ) τd = τimp + τnull + C ( q , q˙ ) q˙ + g ( q )
C2 fimp = Dx (M,J) x˙ + Kx e
C3 fimp = Dx x˙ + Kx e
C4 fimp = Dx x˙ + Kx e + ∫ ki e dt
C5 fimp = Dx x˙ + Kx clip ( e , elim )

Here x tilde = Log(T_ref T_EE^-1), matching the paper's Cartesian pose error.

End effector trajectory

Controller Comparison

Tracking and contact behavior for sampled motion across the implemented controller variants.

Kinesthetic Guidance

Move the simulated arm directly and inspect how human input is converted into constrained motion.

RL Reach Training

A simplified in-browser sparse-reward reach task inspired by the RAM training workflow, with replay, demonstrations, and policy-compatible interventions.

idle waiting for replay intervention inactive

Experimental Setup

Hardware, sensing, and task imagery for the four insertion benchmarks.

Franka robot setup across the four insertion tasks
Franka Emika robot with wrist-mounted RealSense cameras across the task setups.
Robot Franka Emika Robot with Franka Hand end effector
Control compute NVIDIA Jetson Orin AGX running the 1 kHz controller stack
Policy compute Desktop workstation with NVIDIA RTX 4090 for policy and success classifier
Perception Two Intel RealSense D405 wrist cameras, downsampled for policy input

Results

Controller behavior is separated from reinforcement-learning outcomes and task-level metrics.

Controller behavior

Controller tracking error and contact force comparison
Tracking errors are low-pass filtered at 7 Hz for visibility; the closest ablation preserves similar free-space tracking, while simpler controller variants increase tracking error or rely on clipping that causes lag. In the contact test, the proposed stack yields the smallest peak force overall, whereas simpler variants produce substantially higher reaction forces.

Controller Comparison Metrics

Tracking traces are low-pass filtered at 7 Hz for visibility. Peak contact force is measured in the 10 Hz command/check protocol used by the controller comparison.

Controller x error [mm] y error [mm] z error [mm] Peak contact force [N]
C1: PAKT proposed stack16.814.19.622.5
C2: no inertia / desired velocity16.614.19.633.9
C3: manual damping24.822.516.246.5
C4: integral action24.822.416.045.03
C5: clipped error100.537.216.828.15

Learning outcomes

Reward, cycle time, and intervention rate over training
The curves compare the baseline, the hybrid controller variant, and PAKT across reward, cycle time, and intervention rate; RAM insertion reports mean and standard deviation across five seeds, while the other tasks are single real-world runs. Across tasks, PAKT reduces cycle time by 23-48%, demonstration time by 15-34%, and cumulative interventions by 60-85%.

Final Learning Metrics

Baseline is HIL-SERL with SpaceMouse. Hybrid uses the PAKT controller stack with SpaceMouse. PAKT uses the proposed controller stack with kinesthetic guidance. RAM reports five random seeds; other tasks are single real-world runs. Cumulative intervention count is the area under the intervention-rate curve.

Task Baseline: HIL-SERL + SpaceMouse Hybrid: PAKT controller + SpaceMouse PAKT: controller + kinesthetic guidance
Demo time [s] Success rate Cycle time [s] Interventions Demo time [s] Success rate Cycle time [s] Interventions Demo time [s] Success rate Cycle time [s] Interventions
Peg 2161.03.886276 1991.02.252690 1831.02.17980
RAM stick 2430.885.28780 2190.994.075205 1820.984.03325
Limit fixture 2090.993.015937 1940.981.872026 1770.991.921342
Busbar 2210.853.919286 1870.982.092920 1460.992.031327