Nomos research · Satellite incident observability

Something went wrong. Can the satellite tell us why?

We ran four connected ground experiments. One used public satellite telemetry. Two ran on a physical Jetson. One challenged small language models. Together, they test the path from noticing a problem to giving an operator a defensible answer.

Ground evaluationsNo flight claimEvidence downloadable

The four questionsAnomaly → answer
  1. 1NoticeDid behavior change?Measured
  2. 2CaptureWhat caused the failure?Measured
  3. 3PreserveDid the story survive restart?Measured
  4. 4ExplainDid the model follow proof?Gate held

The 60-second answer

The result is simpler than the experiments.

Detecting unusual behavior worked. Explaining it required preserving a small amount of decisive evidence. The language models we tested were not reliable enough to replace that evidence or the operator.

Detection finds the moment.

Public telemetry let a small detector flag unusual segments, but it did not contain enough context to name the cause.

A focused trace keeps the cause.

On the Jetson, small incident records preserved the exact failure and stayed attached to one operation across runtime restarts.

AI should read proof, not invent it.

None of four small language models consistently changed its answer when we removed the decisive evidence.

What this supports: start with tracing for satellite operations. Capture and validate evidence deterministically, then let a model help a human read it.

One question at a time

Four tests. One evidence path.

Each test answers a narrower question than the one before it. Read the conclusion first, then open the workbench if you want the underlying replay.

Public OPS-SAT telemetry

Can we notice unusual behavior?

We trained a transparent detector, then tested it on 529 labelled telemetry segments it had never seen.

See a catch, false alarm, and miss
Useful signal
90/113labelled anomalies found · F1 0.870

The detector also raised four false alerts.

Plain conclusionIt can raise a useful alarm. This dataset cannot tell an operator why the anomaly happened.

Physical Jetson fault injection

Can one symptom hide different causes?

We made an image job fail from memory exhaustion, a frozen worker, and a full disk. From outside the workload, every failure looked identical: the result never arrived.

Replay the three failures
Cause retained
15/15failures explained by the focused record

The largest record was 494 B. At a 4 KiB limit, logs and routine system readings each explained 5/15.

Plain conclusionThe symptom was not enough. A small record saved beside the workload kept the clue that separated the three causes.

Physical Jetson runtime restart

Can the evidence survive a restart?

We followed one imaging operation through two watchdog restarts. The same operation ID and hash-linked journal continued across three runtime boots.

Follow the operation across three boots
Story preserved
25/25failed stages localized at the 2 KiB maximum

Application logs and routine telemetry each localized 20/25. Their five misses were all watchdog-restart cases.

Plain conclusionThe usual downloads showed that the runtime came back. The operation trace kept what it was doing before it disappeared.

Four small open language models

Can a model stay tied to proof?

Each model saw an incident record twice. In the second copy, we removed only the decisive clue. Passing required the model to name the cause with proof, then say “unknown” when that proof disappeared.

Inspect the proof-removal pairs
Safety gate held
0/4models cleared all required pairs

The best result was 6/15 pairs when we supplied the exact proof rules.

Plain conclusionNot yet. A model may help explain verified evidence, but none we tested was reliable enough to diagnose on its own.

The product implication

The trace comes first. The agent comes later.

These results do not point to an autonomous agent flying the spacecraft. They point to a durable operation record that fits beside existing flight software and gives the ground team a coherent story after contact returns.

The agent is not the source of truth. The trace is.

1 · ObserveExisting workload emits events

Command, mode, payload, storage, and downlink boundaries.

2 · PreserveNomos keeps one operation trace

A small durable record survives log rotation and runtime restarts.

3 · ValidateSoftware checks the cited proof

No diagnosis passes unless the record directly supports it.

4 · DecideModel explains. Operator acts.

The human remains responsible for the operational response.

The honest boundary

What we measured. What we still need to learn.

Measured here

  • A held-out public telemetry benchmark
  • Controlled failures on a physical Jetson
  • Persistent traces across runtime-process restarts
  • Deterministic integrity and proof checks

Not proven yet

  • No satellite or commercial flight-software integration
  • No whole-board reboot or power-loss test
  • No private operator telemetry or command history
  • No validated commercial workflow or downlink budget

The next useful test is not another benchmark. It is learning where this record fits inside a real operator’s stack.

Tell us how your team handles incidents

Choose your depth

Start with the replay. Verify anything you want.

The workbench is the fastest way to understand the four tests. The paper and raw files are available for technical review.

Technical narrative

Read the report

Review the protocol, comparisons, system boundary, and citations in one document.

Raw measurements

Verify the evidence

Download measured summaries, model responses, Jetson runs, source snapshots, and hashes.

Ground-bench setup and comparison table

Memory failure

The worker exceeded its 1 GB limit. Linux killed the process.

Frozen worker

The main thread deadlocked. A workload watchdog recorded where it stopped.

Storage failure

The output write hit a full 128 KB test disk and returned “no space left.”

Cause localization at the original 4 KiB download limit
Evidence sentMemoryFrozen workerDisk fullTotal
Application log tailUnknown · 0/5Unknown · 0/5Found · 5/55/15
Routine system readingsFound · 5/5Unknown · 0/5Unknown · 0/55/15
Focused incident recordFound · 5/5Found · 5/5Found · 5/515/15
Complete evidence downloads