Skip to content
AndreaPiPublic

About

My personal home assistant

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Jarvis

Jarvis is a lightweight personal assistant web app. The first module helps you read a water meter photo, review the detected value, and draft an email in Gmail.

Documentation

Features

  • Upload a meter photo and preview it.
  • OCR from a neural-ROI crop with conservative acceptance (unsupported OCR guesses are rejected to manual input).
  • Shadow whole-strip digit reader logging for comparing direct 4-digit predictions against the current per-cell classifier.
  • Auto-fill an email draft with the current date in Italian format.
  • Open a Gmail draft or use a mailto fallback.
  • Run a built-in OCR test set table with Detected, Absolute Error, and Failure Reason columns plus MAE/exact-match/no-read summary stats.

One-click use on macOS

After the backend environment and promoted model files have been installed:

  1. Double-click start-jarvis.command in Finder.
  2. Wait for Jarvis to open in the default browser.
  3. Upload the meter photo, verify or correct the reading, and prepare the email.
  4. Double-click stop-jarvis.command when finished.

The launcher keeps both services on 127.0.0.1, verifies all four promoted model checkpoints through the backend health check, and only stops processes that it started and can identify. If startup fails, the Terminal window stays open with the reason and the local log location.

The same controls are available from a terminal:

npm run jarvis:start
npm run jarvis:status
npm run jarvis:stop

Local Development

  1. Ensure Python 3, uv, and Node.js are installed.
  2. Run the dev server:
npm run serve

Then open http://localhost:8000.

If you also want to run Playwright checks, install JS dependencies once:

npm install

GitHub Push Setup

This repository should push through SSH as AndreaPi. Configure the repo-local remote and SSH key before pushing from a machine with multiple GitHub identities:

git remote set-url origin [email protected]:AndreaPi/Jarvis.git
git config core.sshCommand "ssh -i ~/.ssh/id_ed25519_andreapi -o IdentitiesOnly=yes"
git config user.name "AndreaPi"
git config user.email "[email protected]"

The core.sshCommand setting is intentionally local to this checkout so Git uses the AndreaPi key for this repo without changing other repositories on the same machine.

Optional Neural ROI Backend (recommended)

You can run a Python backend that detects the meter digit window using a fine-tuned pretrained model.

  1. Open a second terminal and set up backend dependencies:
cd backend
uv venv .venv
source .venv/bin/activate
uv pip install -r requirements.txt

Use backend/.venv for any Python workflow that touches images or CV dependencies such as ultralytics, opencv, or Pillow.

For CPU-only environments (for example Vercel), install:

uv pip install -r requirements-cpu.txt
  1. Train/fine-tune a model (copies best checkpoint to backend/models/roi.pt):
python train_roi.py \
  --data data/roi_dataset.yaml \
  --base-model yolov8n.pt \
  --rotation-angles 90,180,270,360 \
  --heavy-augment

The API default ROI checkpoint is pinned to backend/models/roi-rotaug-e30-640.pt. The checkpoint was refreshed on June 9, 2026 after a retrain on the 31-image ROI corpus improved browser OCR MAE from 388.00 to 106.83 with exact match unchanged at 11/31 and no-read unchanged at 1/31. To run with a newly trained checkpoint, set ROI_MODEL_PATH explicitly before starting the backend. train_roi.py now enforces heavy augmentation + rotation expansion by default; weaker runs require explicit --allow-no-augment-policy.

Optional: train the per-cell digit classifier checkpoint:

python train_digit_classifier.py --device cpu

Optional: train the whole-strip shadow reader checkpoint:

python train_strip_digit_reader.py --device cpu

Optional: train the guarded house-specific 23xx strip-reader checkpoint:

python train_strip_digit_reader_23xx.py --device cpu

Thermal monitoring for long macOS training

Use the shared launcher for a long training command. It authenticates for powermetrics in the terminal, waits for the first reported thermal-pressure level (normally up to one 30-second sample), and then starts training under caffeinate:

python3 scripts/train-with-thermal.py -- backend/.venv/bin/python backend/train_full_image_digit_detector.py --fold 0 --device mps

Pass the exact training or resume command after --; the launcher does not change its arguments or checkpoints. It prints a new per-run log directory under backend/runs/training-launcher/. That directory contains training.log, thermal-monitor.log, thermal-events.log for non-nominal samples, and status.json. Training progress is displayed live in the same terminal and retained in training.log; a separate tail is optional. Pass --quiet before -- to hide the training stream while keeping lifecycle and thermal notices. Press Ctrl-C in the launcher's terminal to stop training and monitoring. A terminal hangup (SIGHUP) follows the same shutdown path, including for a wrapped training queue. A failed monitor stops training too; a failed authentication prevents training from starting.

Thermal notices include an hour/minute/second timestamp and appear when pressure changes, including recovery to Nominal. With --auto-pause, notices distinguish macOS thermal state, pause requests, and observed trainer pause/resume. The two thermal sources retain their own labels. Alerts occupy separate lines so progress updates do not overwrite them. Repeated unchanged samples stay in the monitor log.

The training process writes directly to its log file. A separate reader feeds a bounded display queue, and only a dedicated writer touches the terminal. If the output stream is slow or unavailable, live updates may be omitted; thermal control, complete training logs, and process shutdown do not wait for it. Python output is unbuffered; other programs need to flush their own output for immediate display.

Shutdown waits for the owned process groups, including workers whose parent has already exited. If cleanup fails, the launcher exits unsuccessfully and retains the affected group IDs in status.json for inspection. Commands must keep their workers in these groups; detached services are outside this lifecycle.

For a multi-fold Python coordinator, authenticate once for the whole queue:

backend/.venv/bin/python scripts/run-training-queue.py -- path/to/driver.py [driver arguments]

Pass the coordinator file after --, without a second Python executable. It runs in the same process and working directory, preserving its own locking and child-cleanup behavior. The coordinator must wait for its children, stop them on interruption, and launch each training via the updated train-with-thermal.py. Both scripts must include this queue-authentication update. Child environments must inherit JARVIS_TRAINING_AUTH_SESSION (copying os.environ is sufficient).

The queue renews sudo noninteractively every 60 seconds, including evaluation gaps; --renew-interval changes that interval for shorter local sudo timeouts. Each managed fold checks authentication without prompting. A failed renewal blocks subsequent folds; it does not interrupt an already-running monitored training. Queue exit also reports renewal failure. authentication.json under backend/runs/training-queue/<timestamp>/ records authentication status; --log-dir selects a new directory. Renewal stops when the queue exits, without storing passwords, modifying sudo settings, or invalidating other sudo sessions. A stale or missing session fails closed. This does not update an already-running queue: use the wrapper at its next planned start/resume, with the updated launcher.

To run explicitly without thermal monitoring, use --without-monitor before --. This still starts training under caffeinate. The launcher prevents idle sleep while active; it cannot keep a closed MacBook running or guarantee safe hardware temperatures.

Add --auto-pause before -- to enable cooperative pauses for the five backend/train_*.py trainers (full-image digits, ROI, classifier, and both strip readers), using a direct Python invocation from this checkout:

python3 scripts/train-with-thermal.py --auto-pause -- backend/.venv/bin/python backend/train_full_image_digit_detector.py --fold 0 --device mps

The controller reads macOS ProcessInfo.thermalState every second. It requests a pause after 60 seconds of serious, immediately on critical, or when its telemetry is unavailable. The trainer finishes queued accelerator work and waits at the next batch boundary, including during validation. It resumes only after 120 continuous seconds of nominal. Model, gradients, optimizer, scheduler and early-stopping state stay in memory; the pause does not rewrite checkpoints.

A third pause within 30 minutes requires attention and remains paused. Inspect the cause and, when ready, run touch <run-log-directory>/thermal-resume.request; this permits resume only after the same normal-state cooling period. These thresholds are project defaults, not Apple-prescribed safety limits.

thermal-control-events.jsonl records state changes and decisions; thermal-clients/<pid>.json confirms when a trainer actually reaches a pause. A control heartbeat older than 10 seconds, or a missing/invalid control file, blocks the next batch. Loss of the powermetrics logger still stops the run as in monitor-only mode. Ctrl-C also works while paused; restarting a stopped process still requires the trainer's normal checkpoint-resume procedure.

Automatic control supports single-process training. For multi-fold driver scripts, wrap each direct training subprocess separately; unsupported commands are rejected before startup. It cannot be combined with --without-monitor. Without --auto-pause, the launcher's original monitor-only behavior is unchanged. See thermal control details.

The existing monitor also remains available by itself:

scripts/monitor-thermal.sh [--verbose] [log-file]

By default, it shows the first sample and any non-nominal samples in the terminal, recording only complete non-nominal samples in a timestamped log under backend/runs/thermal/. --verbose shows every sample. The launcher's private --no-prompt monitor option uses the authentication obtained before training; standalone use still prompts for sudo normally. The launcher also requests --verbose internally to observe recovery samples, but still stores only non-nominal blocks in thermal-events.log.

Artifact Retention

Treat the following as Tier 1 artifacts that must not be lost:

  • canonical meter photos in assets/
  • assets/meter_readings.csv
  • backend/data/roi_dataset/images/**
  • backend/data/roi_dataset/labels/**
  • backend/data/roi_dataset/splits.json
  • backend/data/full_image_digit_dataset/manifests/**
  • backend/data/digit_dataset/manifests/**
  • backend/data/digit_dataset/sections_synthetic/manifests/**
  • promoted checkpoints in backend/models/*.pt

The repo now uses DVC for the large Tier 1 binaries:

uv pip install --python backend/.venv/bin/python "dvc[s3]"

Currently tracked by DVC:

  • each canonical meter photo in assets/ via per-file *.dvc pointers
  • backend/data/roi_dataset/images via backend/data/roi_dataset/images.dvc
  • backend/data/digit_dataset/windows via backend/data/digit_dataset/windows.dvc
  • backend/data/digit_dataset/windows_canonical via backend/data/digit_dataset/windows_canonical.dvc
  • backend/data/digit_dataset/sections via backend/data/digit_dataset/sections.dvc
  • backend/data/digit_dataset/sections_labeled via backend/data/digit_dataset/sections_labeled.dvc
  • backend/data/digit_dataset/sections_synthetic/train via backend/data/digit_dataset/sections_synthetic/train.dvc
  • promoted model weights in backend/models/*.pt via per-file *.dvc pointers

After dataset ingestion or model promotion:

source backend/.venv/bin/activate
dvc add backend/data/roi_dataset/images
dvc add backend/data/digit_dataset/windows
dvc add backend/data/digit_dataset/windows_canonical
dvc add backend/data/digit_dataset/sections
dvc add backend/data/digit_dataset/sections_labeled
dvc add backend/data/digit_dataset/sections_synthetic/train
dvc add backend/models/*.pt
find assets -maxdepth 1 -type f \( -iname 'meter_*.jpg' -o -iname 'meter_*.jpeg' -o -iname 'meter_*.png' \) -print0 | xargs -0 dvc add
scripts/dvc-push-safe.sh

Do not run raw dvc push directly in this repo. scripts/dvc-push-safe.sh activates backend/.venv, checks that a default DVC remote is configured, and refuses plain local paths and file:// URLs instead of treating them as off-machine storage.

Configure an off-machine DVC remote once before using scripts/dvc-push-safe.sh.

Backblaze B2 is a good default choice for this repo because the current artifact footprint is tiny and fits comfortably within B2's free storage tier. DVC talks to B2 through its S3-compatible endpoint:

source backend/.venv/bin/activate
dvc remote add -d b2 s3://<bucket-name>/jarvis-dvc
dvc remote modify b2 endpointurl https://s3.<region>.backblazeb2.com
dvc remote modify --local b2 access_key_id <key-id>
dvc remote modify --local b2 secret_access_key <application-key>

dvc[s3] must be installed in backend/.venv before this works. The access key and secret should stay in .dvc/config.local, not in committed repo config.

Then create a backup archive when you want a releaseable snapshot:

scripts/package-tier1-artifacts.sh

The generated .sha256 file names the archive by basename, so it can be verified from the download directory with sha256sum -c <archive>.tar.gz.sha256.

For GitHub-hosted retention, set the DVC_REMOTE_URL, DVC_REMOTE_ACCESS_KEY_ID, and DVC_REMOTE_SECRET_ACCESS_KEY repository secrets (plus optional DVC_REMOTE_SESSION_TOKEN if your remote uses temporary credentials), then use the manual Publish Artifacts workflow to dvc pull, package, and upload a release snapshot.

For dataset expansion/QA before retraining:

python plan_digit_expansion.py --target-train-per-digit 12 --priority-digits 4,5,6,9
python validate_digit_dataset.py

validate_digit_dataset.py validates the current windows/canonical/sections workflow.

The replacement full-image digit detector uses a separate, review-first dataset. See docs/full-image-digit-dataset.md for the bootstrap, Make Sense review, guarded import, and image-level cross-validation workflow. Do not train from its labels while any annotation remains pending, and do not use its single historical sanity image for checkpoint selection or as a generalization estimate. Promotion requires a newly collected locked external full-image test set after the recipe is frozen. Evaluate experimental checkpoints with backend/evaluate_full_image_digit_detector.py; complete four-digit exact match and no-read are required alongside YOLO detection metrics. Sources retained only for legacy diagnostics are declared in backend/data/full_image_digit_dataset/manifests/source_exclusions.csv and are omitted from active labels, folds, training, and evaluation. Use npm run qa:full-image-digit-errors after a complete cross-validation recipe to compare its frozen full-image predictions with the production ROI-to-register-crop cascade plus diagnostic register-crop and single-aperture inference, then review the generated visual report before choosing another training recipe. The cascade deliberately predicts one crop at a time to match the deployed endpoint's padding behavior; do not replace it with batched inference when measuring runtime readiness. Use npm run qa:full-image-digit-shadow to run one explicitly configured detector checkpoint as a non-selecting UI shadow. The report compares it with the current OCR on every live test-set image and separately reports the checkpoint's leakage-safe validation-fold slice; only that slice is evidence about unseen active-fold images.

  1. Start the API:
uvicorn app:app --host 127.0.0.1 --port 8001 --reload

In the Codex/DevTools environment, a backend started inside the sandbox may answer shell curl but still be unreachable from the browser. If the page still gets ERR_CONNECTION_REFUSED or Failed to fetch for 127.0.0.1:8001, restart the backend outside the sandbox with escalated permissions and verify from the page context.

By default, the frontend calls http://127.0.0.1:8001/roi/detect and requires neural ROI detection before OCR. Digit decoding is still selected by the per-cell neural classifier at http://127.0.0.1:8001/digit/predict-cells. The whole-strip reader at http://127.0.0.1:8001/digit/predict-strip runs shadow-only and is logged under selectionLog.stripReader. The constrained house-specific reader at http://127.0.0.1:8001/digit/predict-strip-23xx also runs shadow-only and is logged under selectionLog.stripReader23xx; it only accepts a forced 23xx value when its second-digit-is-3 guard reaches the configured threshold. The full-image digit detector endpoint at http://127.0.0.1:8001/digit/predict-full-image-shadow is disabled in the frontend by default. When explicitly enabled, it logs ROI-cropped four-digit candidates under selectionLog.fullImageDigitShadow; the current primary OCR angle selects the headline diagnostic candidate, but the shadow cannot change the selected reading. Check backend readiness with:

curl -s http://127.0.0.1:8001/health

Tests

Run the complete local regression set from the repo root:

npm run test:scripts
npm run test:backend
npm run test:e2e

The script suite covers the one-click launcher safety checks, QA service/checkpoint guards, DVC safety, and artifact packaging. The backend suite auto-discovers backend/test_*.py and covers fast unit, component, and confirmed-regression behavior without training models. Playwright covers browser integration, UI state, neural-ROI failure handling, and OCR selection regressions.

The qa:* commands and the UI Run test set are model/data benchmarks, not additional automated test cases. Add automated coverage only for durable behavior, retained data or artifact safety, a user-facing workflow, or a confirmed regression; consolidate near-identical inputs into a table-driven test.

Generate a per-image ROI checkpoint comparison report (roi-rotaug-e30-640.pt vs roi.pt) with stage 5/6 debug snapshots:

npm run benchmark:roi-diff

This benchmark requires the listed local model files to be present:

  • backend/models/roi-rotaug-e30-640.pt
  • backend/models/roi.pt
  • backend/models/digit_classifier.pt
  • backend/models/digit_strip_reader.pt
  • backend/models/digit_strip_reader_23xx.pt for constrained-reader shadow diagnostics

Report artifacts are written under output/roi-checkpoint-diff/<timestamp>/. Per-image diff tables include selected OCR metadata (sourceLabel, method, preprocessMode) and stage 6 exports use the last 6. OCR input candidate frame from each debug session (the winning decode strip variant).

Generate focused OCR QA artifacts when tuning candidate selection and cell crops:

npm run qa:strip-dataset
npm run qa:ocr-oracle
npm run qa:strip-runtime
npm run qa:cell-crops
npm run qa:roi-geometry-audit
npm run qa:full-image-digit-errors
npm run qa:full-image-digit-shadow
npm run qa:full-image-digit-shadow-sensitivity

These write timestamped reports under output/strip-dataset-qa/, output/ocr-candidate-oracle/, output/strip-runtime-qa/, output/cell-crop-failure-qa/, output/roi-geometry-audit/, output/full-image-digit-error-audit/, output/full-image-digit-shadow-qa/, and output/full-image-digit-shadow-sensitivity/. Use qa:strip-dataset after rebuilding digit windows and before retraining, so the canonical strips can be visually accepted first. The full-image error audit reads frozen evaluation artifacts and never updates canonical annotations. The shadow benchmark identifies the configured checkpoint by path and SHA-256 and distinguishes its complete development comparison from its leakage-safe CV-fold slice. The sensitivity runner uses exact single-image runtime inference and a bounded confidence/NMS grid; its selected fold becomes a tuning surface, not fresh promotion evidence.

The browser-based OCR QA runners fail fast if a reused backend is not ready with the canonical promoted ROI and digit-classifier checkpoints.

CI runs the Chromium Playwright suite on every pull request and on pushes to master. test:scripts and test:backend are currently required local checks and are not run by .github/workflows/e2e.yml.

File Overview

  • index.html: UI layout.
  • styles.css: Styling.
  • app.js: Thin entrypoint that imports src/main.js.
  • src/main.js: UI orchestration and event wiring.
  • src/ocr/: OCR pipeline and neural ROI integration.
  • src/testset/: Manual OCR test-set runner.
  • backend/: Optional FastAPI service for neural ROI and digit-classifier inference/training.
  • AGENTS.md: Repo-wide contributor guide.
  • backend/AGENTS.md: Backend-specific runtime and training guidance.
  • src/ocr/AGENTS.md: OCR-specific behavior, benchmarks, and tuning guidance.
  • assets/: Static assets and example uploads.

Notes

  • OCR now relies on neural ROI detection; if the backend is unavailable or ROI fails, the app asks for manual reading input.
  • Digit decoding uses the backend neural classifier endpoint (/digit/predict-cells) and is enabled by default.
  • The whole-strip digit reader endpoint (/digit/predict-strip) is enabled in shadow mode by default; it logs predictions/debug stage 8 but does not affect the selected reading.
  • The constrained house-specific 23xx endpoint (/digit/predict-strip-23xx) is also shadow-only; it logs accepted/abstained diagnostics and must not affect the selected reading until benchmark evidence supports promotion.
  • The full-image digit endpoint (/digit/predict-full-image-shadow) is optional and frontend-disabled by default; explicit shadow runs log all rotation candidates without changing the primary result.
  • Edge-derived ROI strip candidates are enabled by default and can be toggled with OCR_CONFIG.roiDeterministic.useEdgeCandidates.
  • The selection layer prioritizes edge-derived strips, but the primary classifier pass now also includes top base-strip candidates when they are available; a narrow base fallback rerun is still available only when base candidates were not already evaluated and edge support remains weak. Low-confidence edge-only reads can still be rejected at the final gate.
  • Use the UI Run test set action plus npm run test:e2e for OCR regressions before and after tuning.
  • The Gmail flow opens a draft; you always review and send manually.

Asset Naming (Meter Images)

  • Use the EXIF DateTimeOriginal value as the source of truth for the acquisition date.
  • Rename JPEG/PNG files to meter_yyyymmdd (zero-padded) and keep the original extension.
  • Fully decode HEIC/HEIF imports with the backend Pillow + pillow-heif environment before conversion; metadata-only iCloud placeholders must stop the import without deleting the source. Convert valid inputs to canonical JPEGs named meter_yyyymmdd.JPEG, fully load and verify the JPEG, then delete the original HEIC/HEIF instead of tracking it in DVC.
  • If multiple images share the same date, keep one as-is and add numeric suffixes to the rest (e.g., _1, _2).
  • If EXIF is missing, prefer a known date from the filename or capture notes and document it.

About

My personal home assistant

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages