Nomos-TR-2026-04 · Physical Jetson study

100/100 scene labels. 6/12 priority decisions.

We ran two vision-language models on a physical 8 GB Jetson. The larger one recognized all 100 aerial scenes. Asked to apply a simple priority rule directly, it was wrong half the time.

Physical hardware 200 broad requests 72 repeated task requests Ground bench, not flight
100/100Larger model recognized the scene.
6/12It applied the rule when asked directly.
100/100Simple code applied the rule after recognition.
The experiment

A ground test of the conference setup.

We repeated the conference demo's measurements on the smaller hardware available to us. The model and environment are different.

1

Show the model an overhead image

We preselected 25 public AID images from each of four classes: airport, beach, dense residential, and farmland.

2

Ask three bounded questions

What scene is this? Is it HIGH or ROUTINE under a disclosed rule? What is directly visible?

3

Measure the physical cost

Every request records end-to-end latency, whole-board input power, energy, memory, thermals, exact output, and telemetry errors.

Important: this uses public aerial RGB images on a development board. It is not raw L0 spacecraft imagery, Gemma 3n, a 64 GB AGX, or an in-orbit result.
Test 1 · Perception

The 2.2B model got all 100 labels right.

Four images looked perfect. One hundred images showed what the small model was really doing.

Two confusion matrices. The 500M model gets 46 of 100 images correct and often answers airport. The 2.2B model gets all 100 correct.
One zero-shot pass over 25 preselected images per class. No image was removed. The sample is a calibration set, not the full AID benchmark.
The 500M model did not merely make occasional mistakes. It called every beach an airport. The 2.2B NF4 model separated all four selected classes.
See the outputs

The outputs, exactly as generated.

Choose a model. These are its exact first classification outputs. Greedy decoding returned the same bytes on all three repetitions.

100/100scene correct
4.90 smedian request
10.98 Wmean board power
53.90 Jmean energy
Every parsed label is correct. The period makes every exact label-only contract fail.
Test 2 · Policy

The VLM classified. Code applied the rule.

The policy was deliberately simple: airport or dense residential is HIGH. Beach or farmland is ROUTINE.

Ask the VLM directly

It answered HIGH every time.

AirportHIGH ✓
BeachHIGH ✕
ResidentialHIGH ✓
FarmlandHIGH ✕
6/12correct across three repeats
Parse the scene, then use code

One lookup table. No second model call.

AirportHIGH ✓
BeachROUTINE ✓
ResidentialHIGH ✓
FarmlandROUTINE ✓
100/100derived priorities in the broad sweep
Use the VLM to turn pixels into a bounded observation. Deterministic software can then validate the schema and apply escalation, resource, and action rules.
Test 3 · Physical cost

Accuracy used 2.95× more energy.

The larger quantized package used nearly three times the gross energy per classification.

Accuracy gain
+54 points

46/100 became 100/100 on the same frozen images.

Latency
2.22×

Median end-to-end time rose from 2.21 to 4.90 seconds.

Board power
1.35×

Mean whole-module input power rose from 8.15 to 10.98 watts.

Energy
2.95×

Mean gross request energy rose from 18.28 to 53.90 joules.

Bar charts comparing scene accuracy, latency, board power, and energy for the 500M and 2.2B models.
Physical Jetson Orin Nano in its stock 15 W mode. One hundred measured classification requests per model.
The fit boundary

The BF16 model loaded, then the board rebooted.

The unquantized 2.2B checkpoint technically loaded. The first request never became durable evidence.

2.2B · BF16

Only 377 MiB host memory remained after load. The host rebooted before one request was committed. The watchdog reset source does not prove root cause.

0 requests
2.2B · NF4

Quantization left 2.77 GiB minimum observed host headroom during the broad sweep and completed every image.

100/100
500M · BF16

The small model completed every image with 3.19 GiB minimum observed host headroom. Physical success did not make its labels useful.

100/100 ran
NF4 made the larger checkpoint operational, but its load took 127 seconds. The 500M model loaded in 4.7 seconds. Cold start is part of the design.
Read this honestly

What this does not prove.

These measurements apply to this board, dataset, software stack, and power mode.

Not Gemma

We used pinned SmolVLM2 packages. Model architecture, training, quantization, and prompts differ from the conference result.

Not flight

The Jetson development kit, models, container, and power telemetry are not flight qualified or radiation tested.

Not raw L0

AID is labelled aerial RGB imagery. It does not test sensor ingestion, tiling, radiometric correction, or a real mission distribution.

Not general

One hundred images across four selected classes cannot establish general remote-sensing accuracy. Training contamination was not assessed.

Not an agent

No command generation, flight-software integration, fault containment, digital twin, or human approval workflow was tested.

Open the work

Paper, evidence, and checksums.

The compact public artifacts include exact model revisions, prompts, selected image hashes, physical measurements, failure evidence, and code provenance. Images and model weights are excluded.

Use the model for perception. Keep policy and command authority in deterministic systems.