Adhiraj Singh ← All writing

I Shipped "Edge AI" for a Rover With No Neural Network

TerraSight classifies every patch of planetary terrain a rover sees — compact soil, loose sand, rock, crater, shadow — from the camera feed, on the rover, in real time. The textbook way to do that is a trained neural network. I shipped it without one, and the honest version is better for where the project actually is.

This is the story of building "edge AI" by refusing to train a model until one earns its place.

The rover isn't a cloud GPU

Start with the constraint, because it changes everything. A planetary rover's computer is nothing like the machine you'd train a U-Net on:

So "edge AI" isn't a smaller cloud model. It's whatever fits in that envelope without losing the safety guarantees. And a big labelled dataset of real planetary imagery — the thing a U-Net needs — is exactly what I didn't have.

What a U-Net buys, and whether I needed it

The state of the art is a U-Net: a convolutional encoder-decoder that learns edges → textures → materials from thousands of human-labelled images, and emits a crisp per-pixel class map. It's genuinely the right tool — when you have the labelled data and the compute to train and run it.

I had neither. So I asked the question that decides every one of these: what is the network actually buying me here, versus what can plain feature-space geometry do? For terrain that's largely distinguished by colour and brightness, the answer turned out to be: most of the way there, for none of the cost.

The classical stand-in: nearest colour, honestly

The classifier is a nearest-centroid classifier in a feature space. Each class has a reference "fingerprint"; a pixel is labelled by whichever fingerprint it's closest to.

  1. Features. Turn each RGB pixel into a 6-number vector: (hue, saturation, value, luminance, R−G, G−B). HSV separates colour from lighting; R−G/G−B are cheap opponent-colour channels that pull reddish crater soil apart from greenish mineral edges.
  2. Centroids. Each class's reference colour, run through the same feature function — its fingerprint in the 6-D space.
  3. Classify. Euclidean distance to each fingerprint; nearest wins. Pure-black pixels short-circuit to shadow (unlit, no info); anything matching nothing well becomes unknown.

The clever, safety-relevant part is confidence from margin:

margin = (d2 - d1) / d2                       # d1 = nearest, d2 = runner-up
conf   = min(CONF_CAP, CONF_FLOOR + CONF_CAP * margin)

If one class is clearly closest, confidence is high. If two classes are nearly tied — an ambiguous pixel — confidence drops. And CONF_CAP = 0.6 means a heuristic never claims more than 60% certainty. It's not a trained net; it shouldn't pretend to be. That honesty isn't cosmetic — remember, downstream, low confidence mathematically blocks any "safe" verdict. The classifier being modest is a safety feature.

Quantizing a model that has no weights

Here's where it gets fun. "Shrink the model for the edge" usually means quantizing a neural net's weights. My classifier has no weights — its only float parameters are those class-colour centroids. So I quantized those.

INT8 affine quantization stores each float as an 8-bit integer — 4× smaller, and integer math is far faster and lower-power than float on a GPU-less CPU:

scale      = (max - min) / 255
zero_point = round(-min / scale) - 128
q          = clip(round(x / scale) + zero_point, -128, 127)   # stored int8
dequant(q) = (q - zero_point) * scale                          # recover ≈ x

Quantizing the centroids per-dimension gave ~8× smaller storage. But rounding loses precision, and in a safety system that's the whole risk — so the self-check proves the part that matters: zero class flips across every centroid and every test cell, hazard classes explicitly verified stable, and — critically — confidence only ever decreased. A derate factor guarantees rounding can never make a cell look more confident than it was. Quantization must not manufacture a false-safe, and it's proven not to.

Depth and SLAM, by the way, have no learned parameters at all — so their "shrink" lever isn't quantization but input-resolution and matcher-radius knobs. Same accuracy-versus-cost trade, applied to inputs instead of weights.

Being honest about the numbers

The budget harness times each stage on a fixture: segmentation ~1 ms, depth ~1.3 ms, full run ~5.4 ms, ~45 KiB peak. Those numbers are on a dev laptop, and I refuse to dress them up as rover-readiness — a laptop has no thermal or power throttling. Their real job is to catch a 10–100× regression (an accidental O(n³) loop sneaking into a "measurement" stage), not to certify flight hardware. The real rad-hard envelope lives in the design docs as the target to tune against once real hardware exists. Naming that gap is the point — it's an honest budget framework, not a fake benchmark.

The model isn't cancelled — it's a bounded drop-in

None of this is anti-neural-network. Every stage sits behind a frozen interface, so when a trained model earns its place it drops in without a rewrite. The plan is already written down: MobileNetV3-Small backbone (cheap CPU accuracy, ~2.5 M params, a well-trodden INT8 target), trained → exported to ONNX → post-training INT8 → ONNX Runtime on CPU. And it's gated: the model ships only if it beats the classical baseline on IoU, stays within the latency budget, and keeps the safety regression green. Until it clears those bars, the classical stack is the backbone — not a placeholder for one.

The lesson

"Edge AI" sounds like it must mean a neural network shrunk down small. It doesn't. It means fitting perception into a brutal compute envelope while keeping the safety guarantees — and for terrain that's mostly colour and geometry, a nearest-centroid classifier does that today, with no training data, no GPU, and a confidence model that's honest about being a heuristic.

The neural network is a real upgrade waiting behind a clean seam, with entry criteria it has to earn. Until then, I shipped the whole pipeline without it — because the fanciest component that isn't pulling its weight is just latency, power, and risk you're paying for nothing.