Training robots to fold paper planes! - Attempt 1

Sairaj Loke & Kshitij Bhat · 30th Sept 2026

Compiling our approach + learnings + results for IROS-2026 origmai challenge -https://robotic-origami-challenge.github.io/

Motivation


So the task is simple (ofcourse for humans only yet) - to fold a 15x15 origami Tant paper to a paper airplane. A lot more challenging for a robot given the multi-modal observation dependency, precise manipulation requirements, a lot of occlusion, and contact rich nature of this task. Clearly visible from the videos above ( teleop vs what our best policy could run).

Problem setup


We are provided Sharpa wave teleop data (550+ episodes of 3-4 mins each) for the competition-set. Modalities include RGB-Head-Left/Right, Wrist-Left/Right, tactile(fxfyfzTxTyTz), joint-torques, tactile-images(raw+deformation). Action were 65 dim vectors corresponding to the 65 robot joints.

Things we did


We mainly experimented the capabilities and limitations of existing works name ViTacFormer - https://roboverseorg.github.io/ViTacFormerPage/ and T-Rex - https://tactile-reactive-dexterous.github.io/ robotics policies over this task of folding a origami paper using sharpa robot.

We also did some modifications to the existing architecture in hopes of improving it, including :
1. Handling degenerate action dimensions.
2. Incorporating temporal attention in non-linear window as a proxy for short-term memory (with the intention to handle occlusion).
3. Experimented with different actions prediction techniques (delta actions, absolute actions, hybrid, )
4. Found possible flaws in existing architectures ( shortcut learning in ViTacFormer).

One thing to note is the our validations before testing were done on measuring deviation from actual teleop runs from the validation datasets. Thus doesn't reflect the capability of the trained model but given we didn't have any actual robot / paper folding simulation this was our only resort.

1. ViTacFormer


This was the first policy we adapted to the origami task given it had already worked with sharpa robot albeit on different tasks.

In brief vitacformer itself adpats the pi0 policy to work with extra inputs of tactile.

Limitations we found: It lacks any language encoder hence lacks any ability to incorporate any task prior using language.

Possible flawsin plain architecture:
1. When current state is passed as one of the inputs it clearly performs shortcut learning, or more formally showcases what is known as causal-confusion where the cause of predicted actions is incorrectly deduced from the input actions. Thus is clearly evident in the plots + video below where the policy we tested on real robot clearly does perform paper folding actions even when the paper is not even grasped. Thus associating current state to next state disregarding image modality. We were excited to see the predicts close to ground truth ( which we were also skeptical of) until we saw the real-robot run on - DATE?

This was also evidient in the ablation we did of trained ViTacFormer.Adding large noise in state affects the output MSE by over 400% while noising image only affects by 4% eventho the tokens pertaining to visual domain are lotmore than state tokens. TO ADD: OUR SUBMISSION VIDEOS + PLOTS + ABLATIONS.

We also moved on to T-Rex which we found in later half of August over ViTacFormer given better encoders and overall architecture.

2. T-Rex


Tested on Sharpa robot again, T-Rex is a flow-matching based policy that uses slow KV-cache(involving vision + text tokens) to be attended over by fast queries(of actions/tactiles in fast phase.)

Recent papers like [ list of pprs ]. have moved on to flow-matching approaches given ...FIND THE EXACT REASON......... It was honestly a bit late by the time we started experimenting with T-Rex (already a week into Sept).

Our base policy predicted absolute position we

Runs

REPLACE ME: what you're showing in the run video(s) below.

REPLACE ME: run video caption

Results

REPLACE ME: what the numbers / plots show.

loss curve placeholder
REPLACE ME: loss curve caption
result placeholder 1
REPLACE ME: result 1
result placeholder 2
REPLACE ME: result 2

Conclusion

REPLACE ME: what worked, what the takeaway is.

Mistakes

REPLACE ME: what you'd do differently, what cost you time.

Future plans

REPLACE ME: what's next.