FEnigma/docs/map_vision_plan.md
Dominik Roth c5bcbd3bce docs: map-vision pipeline status + BFS grid-growing plan
Session produced a working prototype (docs/map_vision_wip.py, not
wired into the app yet) for the image-based map-screenshot marker
pipeline: label OCR, pitch estimation from label spacing, a matched-
filter grid-intersection detector, whole-grid crossing prediction, and
RANSAC homography fitting, each validated against real screenshots
with diagnostic images along the way.

Found a real remaining bug before pausing: crossings are currently
all predicted from one single reference label using one global pitch,
so predictions drift with distance from that reference under genuine
perspective distortion (confirmed: ~2-4 degree measured tilt, not
noise). Documented the fix (BFS grid-growing: expand one cell at a
time from every confirmed point, using local rather than global
spacing) as the next step, not yet implemented.
2026-08-09 14:49:40 +02:00

166 lines
9.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Map screenshot vision pipeline: status and next-step plan
Goal: given a screenshot of the in-game map view (not the typewriter text), detect
enemy unit markers (red diamonds per the in-game legend) and resolve each one's
position (large grid + small grid) automatically, for import as targets
(position + id only, type stays `TargetType.UNKNOWN` for this first pass).
This is a *second*, entirely separate OCR/vision pipeline from `ocr.py`'s text
pipeline. `ocr.py`'s module docstring already flags this as future/out-of-scope
work; this doc is that future work's design record.
## Where the code lives right now
`docs/map_vision_wip.py` in this repo is a copy of the working prototype, as of
the end of this exploration session. It is **not wired into the app** and not
`src/fenigma/map_vision.py` yet — it's scratchpad-quality (developed and tested
against three sample screenshots via ad-hoc test scripts, not a proper test
suite). Treat it as a strong starting point, not finished code: variable names,
error handling, and docstrings need a pass before it belongs in `src/`.
## What's validated and working
1. **Grid label OCR** (`find_grid_labels`): sliding-window Tesseract sweep
across the image, filtered to the `[A-T](10|[1-9])` pattern
(e.g. `M8`, `Q10`). Works reasonably well on clean labels; struggles when a
colored hatch-line overlay crosses directly through a label's glyphs
(`L7` misread as `AF`, `N8` as `NWS`) — tried desaturating before OCR,
didn't help (recoloring the hatch line neutral still leaves a shape gap in
the letter). **Needs real inpainting** (fill the interrupted stroke from
surrounding pixels) to fix, not yet done.
2. **Pitch estimation** (`estimate_pitch`): derives cell pitch (both x and y)
from the actual pixel spacing between same-row / same-column labels. No
hardcoded pitch constant anywhere — this was an explicit, correct call-out
mid-session (an earlier version hardcoded `pitch=749` from eyeballing one
image; that doesn't generalize and was thrown out).
3. **Corner/intersection matching** (`find_crossing`, `cross_kernel`): given a
predicted pixel position for a grid-line intersection, finds the real one
nearby. This went through several broken iterations before landing on the
current approach — worth remembering *why* each earlier attempt failed,
so they don't get reinvented:
- Plain local brightness-percentile search: worked on one lucky isolated
line, completely unreliable elsewhere (confirmed via zoomed crop: a
claimed "corner" for `L8` sat in flat background, nowhere near any real
line).
- Generic `cv2.goodFeaturesToTrack`: fires on *any* strong corner, so it
reliably found the label's own text glyphs or a nearby diamond marker's
vertex instead of the grid intersection. Fixed by masking out the
label's own bounding box, then *also* masking out anything with high
HSV saturation (markers and hatching are deliberately colored; the grid
itself is neutral gray/white — a generalizable distinction, not
per-image tuning).
- Even with masking, a generic corner detector still isn't picking the
right *kind* of corner (a line's endpoint kink looks the same as a true
4-way crossing to it). Fixed by replacing it entirely with
`cross_kernel`: a matched filter shaped like a bright `+` (positive
along both a horizontal and vertical arm through the center, negative
in the four quadrant gaps) — a lone single-direction line only lights
up one arm and scores far below a true crossing. This is the piece
that actually made precise matches possible (verified to land within a
couple pixels of a manually-confirmed true corner).
- The matched-filter kernel size and the search-window size must both be
computed **from the current image's own pitch**, never a fixed pixel
constant — a kernel/window tuned for a 750px cell is meaningless on a
150px cell. `find_crossing` takes `pitch_x, pitch_y` and derives both
from them.
- Candidate selection within the window: originally multiplied the
matched-filter response by a tight Gaussian prior (centered on the
predicted position) *before* taking the max — this let a weak,
coincidentally-central false response beat a real, stronger crossing
nearby ("prior is too aggressive, matches predicted center" was the
exact bug report). Fixed by widening the Gaussian (`sigma = 1.5 *
window radius`) and only blending it in lightly (`0.85 + 0.15*prior`)
so the actual filter response does most of the selecting; the prior
now only functions as a mild tie-breaker plus a final plausibility
check on the winner, not the primary selection mechanism.
- The "plausible region" is a Gaussian, so it should be *drawn* as a
circle in diagnostics, not a rectangle — a rectangle visually implies a
hard cutoff that doesn't reflect the actual model. Fixed in the debug
visualization.
4. **Whole-grid crossing prediction** (`span_grid_crossings`): once we have
*any* labels (even 2-3), the grid is regular, so predict and test every
crossing across the visible frame, not just the ones adjacent to a label
that happened to OCR cleanly. This gives far more correspondence points
than "one per successfully-read label." Validated end to end on the
zoomed-out strategic-overview screenshot: 4 labels → 10 predicted
crossings → 4 matched → homography fit with all 4 as RANSAC inliers.
5. **Homography fit** (`fit_grid_homography`): `cv2.findHomography(...,
cv2.RANSAC, 15.0)`, requires 4+ points with real geometric diversity (2+
distinct columns AND 2+ distinct rows — `has_diversity`) before even
attempting a fit, refusing collinear/degenerate input rather than
producing garbage. This replaced an earlier rigid-rotation-only
(translation + single theta) model once it became clear the real screen
has genuine perspective/keystone distortion (parallel lines don't stay
parallel), confirmed by directly measuring the same line's x-position at
two widely-separated y-values and finding a real, consistent ~2-4° drift,
not noise.
## The known remaining bug, and the planned fix
**Bug**: with exactly 4 matched points, `cv2.findHomography` always fits them
*exactly* — 0 residual and "4/4 inliers" is true by construction and doesn't
mean the fit is actually good. Confirmed visually: a 4-point fit on 4 points
that happened to form a lopsided "staircase" in grid-space (missing two
corners of what should've been a proper 2×2 block) produced a visibly
skewed parallelogram instead of a rectangle, even though every individual
point matched its real intersection correctly. The 4 points were individually
right; the *set* was too small and too oddly-shaped to constrain the fit
meaningfully.
Root cause of *why* only 4/10 predicted crossings matched: `span_grid_crossings`
predicts every crossing from **one single reference label** using one global
pitch value for the whole image. Since perspective distortion is real, that
single global (origin, pitch) pair drifts further from the truth the farther a
predicted crossing is from the reference — a jump of several cells accumulates
several cells' worth of drift before the search window even starts looking.
(Some other misses were legitimately unfindable — busy photo texture with no
clean line at that exact spot, confirmed via a zoomed crop — but the long-range
extrapolation drift is the fixable, systematic part.)
### Planned fix: grid-growing (BFS) instead of batch prediction from one origin
Don't predict the whole grid from one fixed point. Walk outward one cell at a
time from every confirmed point, correcting the local estimate as you go:
1. **Seeds**: every label that OCR'd *and* corner-matched successfully is a
confirmed `(grid_col, grid_row) -> (pixel_x, pixel_y)` point. Multiple
seeds, not just the single highest-confidence label.
2. **Expand one cell at a time**: from each confirmed point, only ever
predict its *immediate* neighbor (one cell in one of the 4 directions) —
never extrapolate further than one cell from something already confirmed.
3. **Prefer local spacing over the global average**: if two confirmed points
already share a row (for a column step) or column (for a row step), use
*their* measured spacing to predict the next one out — that's the real
local pitch right there. Only fall back to the global pitch estimate for
the very first step away from a seed, where no local measurement exists
yet.
4. **Search, confirm, repeat**: run the same `find_crossing` search at that
one-cell-away prediction. Success → add to the confirmed set, push onto
the expansion frontier. Failure → that one edge just stops there, doesn't
block expansion from other confirmed points nearby.
5. Keep expanding (a plain BFS/queue over grid coordinates, with a
visited/attempted set so failed edges aren't retried forever) until the
frontier is empty. Feed every confirmed point into the homography fit —
likely dozens of points instead of 4, each individually short-range and
therefore much less exposed to long-range perspective drift, with enough
redundancy that the residual/inlier check from RANSAC actually means
something instead of being a vacuous exact-fit.
This is a genuine restructure of `span_grid_crossings`'s control flow (batch
prediction → BFS), not a parameter tweak. Implementing it is the next step
when this work resumes.
## Also still open (lower priority than the BFS fix)
- Label OCR robustness where hatching crosses the glyph (needs inpainting).
- Once the grid calibration is reliably robust: red-diamond blob detection
(color threshold + connected components) — not started.
- Matching each detected blob to its `#N` id label via nearby OCR — not
started.
- Wiring into `app.py`: a new button, merging results through the existing
`_merge_targets`-style flow as `TargetType.UNKNOWN` — not started.