JAMES WU
DESIGN ENGINEER
02 / Research / SUAT placement / VLA research

From trajectory ranking to force-aware VLA research.

At SUAT’s EMI Lab, I explored how force and motion history could help a manipulation policy anticipate contact risk. I built a trajectory ranker and review GUI, tested simulation routes, then implemented and evaluated an auxiliary risk head during ForceVLA LoRA training.

PROJECTSUAT · Force-Aware VLA Research
INDEPENDENT WORKIndependent research: trajectory ranking, simulation and risk-head training
CONTEXTSUAT · EMI Lab · Independent research project
Mar–Jun 2026
Four stages from trajectory ranking to offline ForceVLA risk-head evaluationInspect full-size image ↗
FIG. 01Research sequence reconstructed from the interim report and placement review. Closed-loop control remains future work.
01 / OVERVIEW

The project in a minute.

THE OUTCOME

A standalone temporal model outperformed a current-force heuristic offline. The integrated ForceVLA risk head also learned the proxy risk signal; reliable action improvement and closed-loop performance remain unverified.

What supports this? ↓
The challenge
Successful manipulation does not describe contact quality. Could recent force-torque and motion history predict near-future instability before an action chunk has finished?
Independent work
Independent research: trajectory ranking, simulation and risk-head training. Standalone AP 0.736 versus 0.571 at h=5; integrated AUROC 0.759 at h=5 and 0.748 at h=15. These are separate experiments with different evaluation settings.
Read this in context
Offline experiments only, without closed-loop robot validation. The standalone GRU and integrated MLP are different models and evaluations; a reliable improvement in action prediction was not established.
Explore the research mapFollow the research stages and results · opens in a new tab↗
02 / THINKING

The choices behind the result.

Open a chapter to follow the work and the choices made along the way.

01 / DESIGN

Make trajectory quality inspectable

I generated Franka trajectory candidates, calculated six cost components and trained a neural ranker. A GUI exposed grasp stability, stress/deformation, smoothness, acceleration, obstacle clearance and path length. This was relative ranking, not a calibrated safety score.

Inspect the supporting evidence ↓
02 / ENGINEERING

Change the route when simulation became impractical

I explored simplified physical proxies, Isaac Lab/TacEx and MuJoCo replay. Contact fidelity, data-generation speed and available compute prevented a credible new dataset within the placement, so I moved to existing ForceVLA trajectories.

Inspect the supporting evidence ↓
03 / ENGINEERING

Test the warning signal before integrating it

A standalone GRU used temporal history from 244 real inputForce trajectories across five tasks. At a five-step horizon, AP was 0.736 versus 0.571 for a current-force heuristic. The reported test set contained 20,475 windows.

Inspect the supporting evidence ↓
04 / ENGINEERING

Separate learning risk from improving actions

I added an auxiliary MLP risk head during ForceVLA LoRA training. Limited offline subsets produced AUROC 0.759 at h=5 and 0.748 at h=15 against force-derived proxy labels. No robust action-loss improvement was established and no closed-loop rollout was conducted.

Inspect the supporting evidence ↓
METHODS USEDTrajectory rankingReview GUIIsaac Lab / TacEx / MuJoCoForceVLA / LoRAGRU / MLP risk headsOffline evaluation
03 / EVIDENCE

Results, and what they mean.

Offline experiments only, without closed-loop robot validation. The standalone GRU and integrated MLP are different models and evaluations; a reliable improvement in action prediction was not established.

0.736 / 0.571

Standalone AP · h=5

Temporal GRU / current-force heuristic. 244 real trajectories, five tasks; 20,475 test windows in the reported study.

0.759 / 0.748

Integrated AUROC · h=5 / h=15

ForceVLA auxiliary MLP head on limited offline subsets with force-derived proxy labels.

Offline

What the results establish

The proxy warning signal was learnable. Safer stopping, retraction, recovery and improved task success have not been demonstrated.

STANDALONE PILOT · AP · h=5

Does temporal history add information?

Same reported pilot; 244 trajectories across five tasks. The integrated model is evaluated separately.

Current-force heuristic0.5710.571
Temporal GRU0.7360.736
Source notes & scope 3
  1. Placement Interim Reportpp. 2–6

    Research context, individual responsibilities and reasons for the simulation-to-ForceVLA pivot.

  2. Placement Review SUAT / Dobot v9slides 3–9

    Later results and explicit limits; read alongside the interim report.

  3. Temporal Force Risk Head — progress (fixed)slide 1

    Standalone GRU / heuristic comparison; integration status on slide 2 predates the later review.

↑ Return to the overview