Studio Notes

Four AI Models Said Our Open Gate Was Closed. So We Trained Our Own.

Local vision models called our open beach gate closed, so we trained a tiny classifier on one scene and let Home Assistant alerts collect the corrections.

Adam Beguelin

By Adam Beguelin

Saturday, September 26, 2026 • Code Puerto Rico, San Juan

Two night camera stills of a wooden double gate with tree-shaped cutouts: closed on the left, ajar on the right.
Closed (left) and ajar (right), same camera and crop, the evening of September 3. The tell is the parting at the center seam. The models were tested on a daytime closed frame and a dusk ajar frame from the same day. Camera stills: Code Puerto Rico

Four of the five local AI vision models we tested side by side looked at our beach gate standing ajar and said it was closed. They gave nearly the same reason: the panels were “parallel and flush.” What runs today is a 6 MB classifier fine-tuned on this one gate and wired into Home Assistant, so every alert on my phone can double as a training label. It still makes mistakes, and I’ll show you those too.

¿Abierto o cerrado?

A camera already watches the gate at a property in Río Grande. I wanted the house to notice when someone leaves the gate open and tell me. The obvious first try was an off-the-shelf vision-language model, an AI that answers questions about images.

On September 3 we ran a quick check with five freely downloadable (open-weight) models through Ollama, which runs models locally, on our GPU workstation (an NVIDIA RTX 4090 we already use for video work). Each got the same crop and the same prompt, not tuned per model, which told them to say open only when the panels were clearly apart. They saw two frames: the gate closed in daylight and ajar at dusk. Four of them (qwen3.5:9b, qwen3.8:27b, qwen3-vl:4b and mistral-small3.1) got the closed frame right and called the ajar one closed. The fifth, moondream:1.8b, gave no usable answer.

One very large cloud model, qwen3.5:397b on Ollama Cloud, did catch the ajar frame. We ran it in production for about a day, asking it three times per check and going with the majority. Two frames is not a benchmark, and we never tested that model on our full photo set. Still, I wanted the decision made inside the house, not by three cloud calls every check.

A small model for one scene

So we trained a classifier. It is MobileNetV3-Small, a compact image network pretrained on ImageNet, fine-tuned in PyTorch to answer a single question. We chose it because it is small enough to run without a GPU. A first version went live September 5. The one we ran until September 25 was retrained September 6 on 108 labeled frames and got 21 of 22 held-out images right. Held-out images are photos kept out of training to test the model. Only 6 of those 22 show the gate open, so the score says less than it seems.

Training happens on the 4090: the September 25 run took 16.5 seconds wall-clock (40 passes over 120 photos, start-up included). The day-to-day checks run on a small CPU-only server; the workstation is needed elsewhere. Every five minutes the server grabs a still, crops it to the gate and classifies it. Each check takes about six seconds, model loading included.

Diagram: camera to classifier to Home Assistant, which sends a reminder and a phone push; a vote on the phone feeds the training set, which feeds a retrain on the GPU, which sends an updated model back to the classifier.
The top row runs on its own every five minutes. The orange path is the human in the loop, and it is what makes the model better. Diagram: Code Puerto Rico

Home Assistant closes the loop

The classifier’s verdict flips a switch in Home Assistant, the open-source platform that already runs our lights, AC and cameras. If the gate reads open for 15 minutes, Home Assistant sends a reminder.

iPhone notification titled Gate Still Open Alert: Beach Gate has been open for 45 minutes, with the camera view of the open gate and three buttons: Gate is open, Gate is closed, and Snooze until tomorrow.
The reminder when the gate stays open: 45 minutes in, with a third button to snooze until tomorrow. Screenshot: Code Puerto Rico

Each time the verdict changes, my phone gets a push with the photo and two buttons: “Gate is open” and “Gate is closed.” A tap saves that frame, with my label, for the next retrain. If the gate flips to closed and I tap “Gate is open,” the frame becomes an open example and the alerts keep coming until I tap “Gate is closed.” If I ignore an alert, the verdict stands and the alerts stop. As of September 25, a vote only counts if it arrives within about 20 minutes of the photo. The catch: I can only correct the model when it changes its mind. A gate it wrongly calls closed sends no alert, as I found out.

iPhone notification: Beach gate CLOSED, all clear, with a camera thumbnail showing the gate and two buttons, Gate is open and Gate is closed.
A real alert. Tapping either button files the photo as training data; ignoring it lets the verdict stand. “conf” is the model’s score for its verdict, not a calibrated probability. Snapshot link and background redacted. Screenshot: Code Puerto Rico
iPhone notification titled Beach gate OPEN, showing the camera view of the beach gate and two buttons: Gate is open and Gate is closed.
The real alert on my phone: the classifier’s guess, the photo, and two buttons to label it. Screenshot: Code Puerto Rico

The miss, and my own bad labels

On September 25 I left the gate open for about 18 minutes one afternoon and got no alert. In low afternoon sun and hard shadows, the model gave three open frames an open score (0 to 1, where 0.5 or more means open) of 0.13, 0.001 and 0.001. That is a miss, or false negative. The opposite error, calling a closed gate open, is a false alarm (false positive).

Reviewing every labeled photo that day turned up a second problem: me. Three frames I had tapped “Gate is closed” clearly show the gate open. The model had called all three open, with scores of 0.98 to 0.99. My taps came seven hours to four days later. I was probably describing the gate as it was when I tapped, or reading “Gate is closed” as “I closed it.” One of those labels went into the September 6 retrain, and that model learned my mistake: it scored that wide-open frame 0.00.

My labeling page shows each alert frame with the model’s guess, my label, and Open and Closed buttons to change it.

Tall phone screenshot of a dark-mode web page titled Beach gate: 29 open alerts, 0 still unlabeled, with links for Alerts sent, Unlabeled and All frames. Below are eight cards from September 26 back to September 20, day and night, each with a gate camera frame, a timestamp, a pred open badge with a confidence percentage, a labeled open or labeled closed badge, and green Open and red Closed buttons.
The top of my labeling page: each alert frame, the model’s guess and confidence, my label, and two buttons to fix it. In the last two cards, the model said open and my label said closed. Screenshot: Code Puerto Rico

Then a surprise: rerun on those three frames, our first model, from September 5, scored them 0.99, 0.55 and 0.55. Did my bad label break the retrain? I can’t show that. Retraining on the same photos with different random starting points swung the scores from near 0 to 0.8, and removing the label didn’t reliably help. One training run proves little.

Six gate frames. Top row: three open gates I mislabeled closed. Bottom row: a glare-streaked closed gate the model called open, and two partly open gates in afternoon sun it called closed.
Top: the model was right and my labels were wrong. Bottom: real model errors, a 4:30 am glare streak and two of the September 25 afternoon misses. Camera stills: Code Puerto Rico

We fixed the three labels and added 11 hand-checked frames, for 150 labeled photos in all (63 open, 87 closed). We also used augmentation, adding altered copies of training photos with fake sun glare, hard shadows and tilted angles. Then we retrained. The table compares the two models, rerun from their saved files.

What we testedSep 6 modelSep 25 model
The original 22 held-out images21 of 22 right21 of 22 right
The 3 afternoon misses0 of 3 caught3 of 3 (it trained on them)
The same 3, model trained without them–1 of 3 caught
False alarms, 789 frames believed closed (Sep 23–25)00

In production, scores under 0.6 are marked unknown; that changes no result here.

Held-out accuracy didn’t move, and both models still miss a night frame with a person standing in the open gateway. The retrained model catches the September 25 misses mostly because it has seen them; with them held out, it caught one of three. Our test scores also flatter the model: some test photos are near-twins of training photos taken minutes apart. The fix is more afternoon examples, which the phone votes can now supply.

What anyone using AI can take from this

  • Test before you trust. A two-frame test told us more than any model card.
  • Go small and specific when the camera never moves and the question never changes.
  • Audit your labels. Ours had errors, and the button wording invited them. Next step: “Photo shows open / closed.”
  • Hold out your failures, and retrain more than once, or you’re grading your own homework.
  • Connect the model to something real. Home Assistant is what turns a score into a reminder, and a reminder into new training data.

Test, measure, fix the data, ship, repeat: that loop is applied AI, and it’s the habit we build at Holberton Coding School Puerto Rico, Code Puerto Rico’s school, through its AI Software Engineering program and the part-time AI for Developers program for working developers. Ask us about the gate at the Holberton Coding School Puerto Rico booth at the Caribbean AI Summit, October 9–10 at the Puerto Rico Convention Center.

Follow Code Puerto Rico on Instagram: @code_puertorico