Show the model an overhead image
We preselected 25 public AID images from each of four classes: airport, beach, dense residential, and farmland.
We ran two vision-language models on a physical 8 GB Jetson. The larger one recognized all 100 aerial scenes. Asked to apply a simple priority rule directly, it was wrong half the time.
We repeated the conference demo's measurements on the smaller hardware available to us. The model and environment are different.
We preselected 25 public AID images from each of four classes: airport, beach, dense residential, and farmland.
What scene is this? Is it HIGH or ROUTINE under a disclosed rule? What is directly visible?
Every request records end-to-end latency, whole-board input power, energy, memory, thermals, exact output, and telemetry errors.
Four images looked perfect. One hundred images showed what the small model was really doing.
Choose a model. These are its exact first classification outputs. Greedy decoding returned the same bytes on all three repetitions.
The policy was deliberately simple: airport or dense residential is HIGH. Beach or farmland is ROUTINE.
HIGH ✓HIGH ✕HIGH ✓HIGH ✕HIGH ✓ROUTINE ✓HIGH ✓ROUTINE ✓The larger quantized package used nearly three times the gross energy per classification.
46/100 became 100/100 on the same frozen images.
Median end-to-end time rose from 2.21 to 4.90 seconds.
Mean whole-module input power rose from 8.15 to 10.98 watts.
Mean gross request energy rose from 18.28 to 53.90 joules.
The unquantized 2.2B checkpoint technically loaded. The first request never became durable evidence.
Only 377 MiB host memory remained after load. The host rebooted before one request was committed. The watchdog reset source does not prove root cause.
Quantization left 2.77 GiB minimum observed host headroom during the broad sweep and completed every image.
The small model completed every image with 3.19 GiB minimum observed host headroom. Physical success did not make its labels useful.
These measurements apply to this board, dataset, software stack, and power mode.
We used pinned SmolVLM2 packages. Model architecture, training, quantization, and prompts differ from the conference result.
The Jetson development kit, models, container, and power telemetry are not flight qualified or radiation tested.
AID is labelled aerial RGB imagery. It does not test sensor ingestion, tiling, radiometric correction, or a real mission distribution.
One hundred images across four selected classes cannot establish general remote-sensing accuracy. Training contamination was not assessed.
No command generation, flight-software integration, fault containment, digital twin, or human approval workflow was tested.
The compact public artifacts include exact model revisions, prompts, selected image hashes, physical measurements, failure evidence, and code provenance. Images and model weights are excluded.
Use the model for perception. Keep policy and command authority in deterministic systems.