Release evaluation for robotics teams

Know whether your robot policy is ready to ship.

Sentinuum collects test results from simulation, hardware runs, and deployed fleets, so your team can compare policy versions, investigate failures, and record release decisions.

01 / THE PROBLEM

Robot policy evaluation is hard to scale.

As policies take on more tasks, teams need to compare evidence from simulation, recorded runs, and hardware tests. The methods and metrics often differ across those environments, which makes release decisions harder to reproduce. Sentinuum gives the team one traceable record from each episode to the decision to ship, hold, or test again.

SIMULATIONRECORDED LOGSHARDWARE TESTSVERSIONED METRICS

02 / PRODUCT WORKFLOW

From test run to release decision.

01

Observe

Collect outcome-bearing episodes from simulation, bench tests, and fleet runs.

02

Evaluate

Score episodes against a versioned evaluation contract.

03

Investigate

Trace regressions and group repeated failure modes.

04

Decide

Record the evidence, threshold, owner, and release state.

03 / CAPABILITIES

What the team can do.

Sentinuum keeps test results, policy versions, failure analysis, and release criteria connected so engineers can review what changed and decide what to do next.

01

Normalized episode evidence

One evidence model across simulation, bench tests, and fleet telemetry.

02

Versioned evaluation contracts

Pin scenarios, metrics, thresholds, and policy versions to every decision.

03

Failure clustering

Find repeated failure modes without reducing every run to a pass/fail count.

04

Simulator-to-hardware comparison

See where simulated results hold on hardware and where they diverge.

05

Regression history

See what changed, when it changed, and which episode groups moved.

06

Release gates

Set measurable ship criteria and review them before each release.

07

CI & webhook exports

Move release state into the engineering systems your team already uses.

08

Prioritized next actions

Focus the next test run on the evidence gap most likely to change a decision.

04 / ARCHITECTURE

Your systems stay yours.

Your robots, simulators, telemetry, and source data stay with you. Sentinuum runs the normalization, evidence lineage, evaluation, and release workflow on top.

CUSTOMER-OWNED
RobotsSimulatorsTelemetrySource data
EVENTS + ARTIFACTS
SENTINUUM PLATFORM
Normalization
Evidence lineage
Evaluation engine
Release workflow

TEXT-BASED ADAPTERS

MuJoCoIsaac SimGazeboROS 2GitHubGitLabInternal stacks

05 / CURRENT PROOF

463

REAL ROBOT-MANIPULATION EPISODES

A reference prototype, measured on public data.

The current prototype uses 463 real robot-manipulation episodes from the UCI Robot Execution Failures dataset (CC BY 4.0). Every metric is measured or deterministically derived from the underlying evidence.

This is a reference prototype—not a customer production deployment.

06 / WHO WE BUILD WITH

01

Warehouse & logistics robotics

02

Industrial manipulation

03

Humanoids

04

Autonomous mobile robots

05

Drones

06

Medical & assistive robotics

07 / DESIGN PARTNERS

Start with one release decision.

In a six-week paid pilot, we connect one robot program and one data source, define the release criteria with your team, and evaluate an active policy decision.

Discuss a pilot
Six-week pilot01
One robot program02
One telemetry adapter03
One active release decision04
Measurable success criteria05

08 / COMPANY

Built for the teams shipping physical AI.

Sentinuum is built by a technical founder with extensive experience across physical AI and platform engineering, including work in the Stanford ecosystem. That background spans robotics evaluation, production software, and the infrastructure needed to turn complex test data into decisions engineers can act on.

We are currently working with robotics teams on active policy-release decisions and looking for design partners who want a more rigorous, repeatable evaluation process.

CONTACT

Talk with us about a release decision.

We use this information only to respond to your inquiry.

09 / FAQ

Direct answers.

Does Sentinuum replace simulators?

No. Simulators still generate the test data. Sentinuum makes the results comparable and traceable.

Does Sentinuum certify safety?

No. Sentinuum supports evidence-based release decisions; it is not a safety certification body.

Who owns the data?

The customer retains ownership of source data. Sentinuum owns its platform, normalization, and evaluation logic.

What is required to onboard?

An active release, at least 200 outcome-bearing episodes, and a technical owner.