Engineering Blog

An Update on the Sidekick Flywheel

The Sidekick Flywheel nearly tripled task success on towel folding, lifting the success rate from 18.0% to 51.5%, a gain of 33.5 percentage points on one of the hardest problems in physical manipulation.

Emailfounders@sidekickrobotics.ai
The Sidekick Flywheel — towel folding success results

Towel-folding success with and without the Sidekick Flywheel. Both arms ran the same system on the same task under the same conditions, with the flywheel as the only variable. Bars show 95% Wilson intervals over 200 episodes without the flywheel and 99 episodes with it.

Robots are good at objects that hold their shape. A bolt is a bolt every time you pick it up, in the same place, with the same edges. That predictability is what makes industrial automation work, and it is also what confines it — to fixtured parts, controlled lighting, and environments built around the machine rather than the person.

Cloth breaks all of it. A towel has no canonical pose. It bunches, drapes, slides against itself, and changes shape in response to the very contact you use to manipulate it. There is no single correct grasp, and the target keeps moving while you reach for it. Tasks like this are where general-purpose robotics has to succeed if it is going to leave the factory floor.

Today we are sharing where the Sidekick Flywheel has taken us on that problem.

2.9×

the success rate of the same system without the flywheel

+33.5

percentage point gain, 18.0% → 51.5%

299

episodes in the comparison, every one published

Nearly tripling success on a deformable-object task

The headline number is a straightforward one: on towel folding, the flywheel takes the system from succeeding roughly one attempt in six to succeeding better than one in two. In practical terms that is the difference between a demonstration and a capability, between something that occasionally works when conditions are kind, and something that finishes the job more often than not.

What makes the result meaningful is the shape of the comparison behind it. Both arms are the same system, running the same task, under the same conditions. One had the flywheel and one did not, and nothing else was permitted to vary. That design is what lets us attribute the entire gap to the flywheel rather than to any of the many other factors that move a robot's success rate around.

Success rate comparison: 18.0% without flywheel vs 51.5% with flywheel

The two arms side by side. Success rises from 18.0% to 51.5%, an absolute gain of 33.5 percentage points. Error bars show the 95% confidence interval on each rate — the range within which the true success rate is expected to fall.

Every attempt on the record

Aggregate percentages ask the reader for a measure of trust. We would rather supply the evidence instead. Below is every single attempt in the comparison: 200 without the flywheel and 99 with it. Each dot is one try at the towel, filled where the fold succeeded.

Publishing at this resolution is a deliberate standard for us. The denominator travels with the number, the failures are as visible as the successes, and anyone can count the difference for themselves.

All 299 episodes — each circle is a single attempt, filled circles are successful folds

All 299 episodes in the comparison. Each circle is a single attempt; filled circles are successful folds. Nothing has been averaged, excluded, or summarised away.

Measured with the uncertainty attached

Every measurement taken on a finite number of attempts carries uncertainty with it, and we think the honest practice is to publish the size of that uncertainty rather than the single most flattering number inside it. Across the full range of gains consistent with our evidence, the flywheel's contribution stays large and stays positive.

Distribution of the true gain — centred at +33.5 percentage points, 95% interval +22.1 to +44.2

The distribution of the true gain implied by the evidence, centred at +33.5 percentage points with a 95% interval from +22.1 to +44.2. The entire distribution sits above zero.

The finding holds under whichever standard measure you prefer to read it in.

Result expressed as absolute gain, multiple of base rate, and odds ratio with 95% confidence intervals

The same result expressed as an absolute gain, a multiple of the base rate, and an odds ratio, each with its 95% confidence interval. In all three the no-effect value falls well outside the interval.

Building toward general-purpose physical work

Towel folding is a test, not a product. We chose it because it is honest about the things that make real environments hard: deformable material, no fixed target, and a task where success is unmistakable to anyone watching. A system that improves here is improving at the properties that generalise, not at a single memorised motion.

That is the direction of travel. The useful property of a flywheel is that it is indifferent to what it is pointed at, and the work ahead is pointing it at progressively more of the world: more materials, more tasks, more of the ordinary physical work that has stayed out of reach of automation because it refuses to hold still.

We will keep publishing what we measure, at this resolution, as we build the AI Platform for Physical Work. Any model. Any robot. Any task. Reliable general purpose AI for the physical world.

About these results

Comparison conducted on a towel-folding task under fixed conditions. 36 of 200 episodes succeeded without the Sidekick Flywheel; 51 of 99 succeeded with it. Confidence intervals are 95% Wilson. Absolute difference: +33.5 percentage points, 95% CI +22.1 to +44.2.