Reads a screenshot of the in-game map table and recovers where on the board it is looking. The scene is a flat table under a perspective camera, so a single homography describes grid<->screen exactly; that is fitted from line evidence (LSD segments, vanishing-point RANSAC, 1-D lattice fits). Line evidence alone cannot finish the job: the 1km cells and the 100m subgrid are locally identical, so it fixes neither scale, axis assignment, axis direction nor phase. Those come from the cell labels -- and the labels are never *detected*, they are correlated where the grid says they must be (9% in from a cell's left edge, 6% down from its top). Blob detection on aerial-photo terrain finds texture, not glyphs; correlating a known template at a known place has no such failure mode. The same score then ranks the geometry hypotheses, since a wrong lattice puts the crop where no label is, so one number resolves scale, axis assignment, direction, phase and anchor together. Markers are found by hostile/friendly colour plus an IoU test against an ideal inscribed diamond, which cut a red-lit shot from 43 false positives to 2 while preserving every hand-verified count. Unit-type classification is present but not yet reliable, and returns unknown rather than guessing. Measured over the fixture set: 7 of 10 solve, each with 100% of its ground-truth points in the correct cell (85/112 overall), residual spread 0.005-0.033 cells. The other three reject cleanly; none has ever produced a plausible-but-wrong grid. Across 122 typewriter screenshots solve() accepted none, which is what makes clipboard routing safe. Fixtures: 10 map shots with hand-transcribed ground truth, plus 10 typewriter shots spanning 262-5366px for routing and future OCR tests. Map shots wider than 2400px (the pipeline's own maximum working width) were downscaled with their coordinates rescaled to match; re-measured afterwards, the results are identical. Typewriter shots stay at native resolution because OCR needs the text legible. Drops docs/map_vision_plan.md and its WIP prototype: the design now lives in the module docstring, next to the code it describes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
909 lines
37 KiB
Python
909 lines
37 KiB
Python
"""Recover the game grid from a screenshot of the map table, and read the
|
|
unit markers off it.
|
|
|
|
This is a second, separate pipeline from `ocr.py`'s typewriter-text OCR.
|
|
See `docs/map_vision_plan.md` for the design record, the measurements
|
|
behind it, and the approaches that were tried and rejected.
|
|
|
|
The scene is a *flat* table viewed by a perspective camera, so a single
|
|
homography describes grid-to-screen exactly. The grid, its per-cell labels
|
|
and the markers are drawn on the table surface, so they are visible whether
|
|
or not the aerial photo covers that part of the map.
|
|
|
|
Shape of the solution:
|
|
|
|
1. Line evidence -> two projective line families -> candidate lattices.
|
|
Abundant and cheap, but it can only ever give a lattice up to unknown
|
|
scale (1 km cells and the 100 m subgrid look identical), axis
|
|
assignment, direction and phase.
|
|
2. Those discrete unknowns, plus the absolute anchor, are resolved by
|
|
READING CELL LABELS -- and the labels are not detected. Once a
|
|
candidate lattice exists we know exactly where a label must be (9% of
|
|
a cell in from the left, 6% down from the top), so we crop there and
|
|
correlate the glyphs rendered in the game's own font. Detection was
|
|
tried five different ways and always returned aerial-photo texture
|
|
instead of glyphs; correlating a known template at a known place does
|
|
not have that failure mode.
|
|
3. The correlation score also *ranks the lattice candidates*: a wrong
|
|
lattice puts the crop where no label is, so it scores low. One number
|
|
therefore selects scale, axis assignment, direction, phase and anchor
|
|
together.
|
|
|
|
Measured on the 10 fixtures in tests/fixtures/map_shots: solves 7 of them,
|
|
with 100% of each solved shot's annotated points landing in the correct cell
|
|
(85 of 112 overall) and a residual spread of 0.005-0.033 cells. The other
|
|
three are rejected rather than guessed at, and no fixture has ever produced
|
|
a plausible-but-wrong grid. Rejection is a supported outcome -- a silently
|
|
misplaced target is far worse than a refusal.
|
|
|
|
Over the 122 typewriter screenshots this was checked against, solve()
|
|
accepted none, which is what makes it safe to route clipboard images through
|
|
it (see looks_like_map for the cheap pre-filter).
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import numpy as np
|
|
from pathlib import Path
|
|
|
|
try:
|
|
import cv2
|
|
except ImportError as exc: # pragma: no cover
|
|
raise ImportError(
|
|
"map_vision needs opencv (pip install opencv-python-headless)") from exc
|
|
|
|
from PIL import Image, ImageDraw, ImageFont
|
|
|
|
LARGE_X = "ABCDEFGHIJKLMNOPQRST"
|
|
COLS, ROWS = 20, 10
|
|
FONT_PATH = Path(__file__).resolve().parents[2] / "assets" / "fonts" / "CourierPrime-Regular.ttf"
|
|
|
|
WORK_W = 1500 # working resolution; screenshots vary 700..6880 px wide
|
|
CELL_PX = 240 # canonical size a rectified cell is warped to
|
|
CELL_MARGIN = 0.18 # rectify beyond the cell bounds, see rectify_cell()
|
|
PAD_L, PAD_T = 0.09, 0.06 # label padding inside its cell (game constant)
|
|
LABEL_ACCEPT = 0.62 # per-read confidence; measured: correct reads 0.73-0.87,
|
|
# wrong reads 0.40-0.56, so this sits inside the gap
|
|
MIN_LABEL_VOTES = 2 # one label has no error detection: a misread shifts the
|
|
# whole board with nothing to contradict it
|
|
FLIPS = ((False, False), (True, False), (False, True), (True, True))
|
|
SWAP = np.array([[0, 1, 0], [1, 0, 0], [0, 0, 1]], np.float64)
|
|
LABELS = [f"{c}{n}" for c in LARGE_X for n in range(1, 11)]
|
|
|
|
|
|
# --------------------------------------------------------------- evidence
|
|
|
|
def load(path, work_w=None) -> np.ndarray:
|
|
img = cv2.imread(str(path), cv2.IMREAD_COLOR)
|
|
if img is None:
|
|
raise ValueError(f"cannot read image: {path}")
|
|
return downscale(img, work_w)
|
|
|
|
|
|
def downscale(img, work_w=None) -> np.ndarray:
|
|
h, w = img.shape[:2]
|
|
s = min(1.0, (work_w or WORK_W) / w)
|
|
if s < 1.0:
|
|
img = cv2.resize(img, (int(w * s), int(h * s)), interpolation=cv2.INTER_AREA)
|
|
return img
|
|
|
|
|
|
def ridge(img: np.ndarray) -> np.ndarray:
|
|
"""Bright, neutral, thin line structures; coloured overlays suppressed.
|
|
|
|
Two things this must get right:
|
|
|
|
* the colour mask is relative to the scene's OWN illuminant. One fixture
|
|
is lit bright red, and an absolute saturation cut masks the entire
|
|
table, grid included.
|
|
* several top-hat kernel sizes, not one. A top-hat kernel must be LARGER
|
|
than the structure it keeps or it hollows it out, and grid lines run
|
|
from ~2px when the whole table is in frame to ~15px when a single cell
|
|
fills it.
|
|
"""
|
|
lab = cv2.cvtColor(img, cv2.COLOR_BGR2LAB)
|
|
lum = lab[:, :, 0]
|
|
a = lab[:, :, 1].astype(np.float32) - float(np.median(lab[:, :, 1]))
|
|
b = lab[:, :, 2].astype(np.float32) - float(np.median(lab[:, :, 2]))
|
|
chroma = np.sqrt(a * a + b * b)
|
|
acc = np.zeros(lum.shape, np.float32)
|
|
for k in (7, 15, 31, 51):
|
|
el = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (k, k))
|
|
acc = np.maximum(acc, cv2.morphologyEx(lum, cv2.MORPH_TOPHAT, el).astype(np.float32))
|
|
acc[chroma > 16.0] = 0
|
|
return (np.clip(acc, 0, 55) / 55.0 * 255).astype(np.uint8)
|
|
|
|
|
|
def segments(ev: np.ndarray, min_len: float) -> np.ndarray:
|
|
lines = cv2.createLineSegmentDetector().detect(ev)[0]
|
|
if lines is None:
|
|
return np.zeros((0, 4), np.float32)
|
|
L = lines.reshape(-1, 4)
|
|
return L[np.hypot(L[:, 2] - L[:, 0], L[:, 3] - L[:, 1]) >= min_len]
|
|
|
|
|
|
def _homog_lines(L):
|
|
p1 = np.c_[L[:, 0], L[:, 1], np.ones(len(L))]
|
|
p2 = np.c_[L[:, 2], L[:, 3], np.ones(len(L))]
|
|
ln = np.cross(p1, p2)
|
|
return ln / (np.linalg.norm(ln[:, :2], axis=1, keepdims=True) + 1e-9)
|
|
|
|
|
|
def _vp_ransac(L, allowed, iters=4000, tol_deg=1.5, seed=0):
|
|
idx = np.flatnonzero(allowed)
|
|
if len(idx) < 3:
|
|
return np.zeros(len(L), bool)
|
|
rng = np.random.default_rng(seed)
|
|
ln = _homog_lines(L)
|
|
mid = np.c_[(L[:, 0] + L[:, 2]) / 2, (L[:, 1] + L[:, 3]) / 2]
|
|
ang = np.arctan2(L[:, 3] - L[:, 1], L[:, 2] - L[:, 0])
|
|
tol = np.deg2rad(tol_deg)
|
|
best = np.zeros(len(L), bool)
|
|
for _ in range(iters):
|
|
i, j = rng.choice(idx, 2, replace=False)
|
|
v = np.cross(ln[i], ln[j])
|
|
if abs(v[2]) < 1e-9:
|
|
continue
|
|
vx, vy = v[0] / v[2], v[1] / v[2]
|
|
want = np.arctan2(vy - mid[:, 1], vx - mid[:, 0])
|
|
diff = np.abs((want - ang + np.pi / 2) % np.pi - np.pi / 2)
|
|
inl = (diff < tol) & allowed
|
|
if inl.sum() > best.sum():
|
|
best = inl
|
|
return best
|
|
|
|
|
|
def two_families(L, min_sep_deg=20.0):
|
|
"""The two pencils of grid lines.
|
|
|
|
The second family is required to be angularly DISTINCT from the first.
|
|
Simply re-running RANSAC on the leftovers lets both passes lock onto the
|
|
same family and report two "families" a degree apart. Orientation is
|
|
also undirected, so angles are compared as doubled angles -- otherwise
|
|
+89 and -89 degrees look like opposites instead of neighbours.
|
|
"""
|
|
th2 = 2.0 * np.arctan2(L[:, 3] - L[:, 1], L[:, 2] - L[:, 0])
|
|
inA = _vp_ransac(L, np.ones(len(L), bool), seed=1)
|
|
if inA.sum() == 0:
|
|
return inA, np.zeros(len(L), bool)
|
|
mA = np.arctan2(np.median(np.sin(th2[inA])), np.median(np.cos(th2[inA])))
|
|
d = np.abs(np.angle(np.exp(1j * (th2 - mA)))) / 2.0
|
|
inB = _vp_ransac(L, (~inA) & (d > np.deg2rad(min_sep_deg)), seed=2)
|
|
return inA, inB
|
|
|
|
|
|
def _fit_vp(L, inl):
|
|
_, _, Vt = np.linalg.svd(_homog_lines(L[inl]))
|
|
v = Vt[-1]
|
|
return v / (v[2] if abs(v[2]) > 1e-12 else 1e-12)
|
|
|
|
|
|
# ---------------------------------------------------------- lattice fitting
|
|
|
|
def _dir_of(Hp, sub):
|
|
ds = []
|
|
for (x1, y1, x2, y2) in sub:
|
|
p = Hp @ np.array([[x1, x2], [y1, y2], [1.0, 1.0]])
|
|
if np.any(np.abs(p[2]) < 1e-9):
|
|
continue
|
|
p = p[:2] / p[2]
|
|
d = p[:, 1] - p[:, 0]
|
|
n = np.linalg.norm(d)
|
|
if n > 1e-9:
|
|
d = d / n
|
|
ds.append(d if d[0] >= 0 else -d)
|
|
if not ds:
|
|
return None
|
|
d = np.median(np.array(ds), axis=0)
|
|
return d / (np.linalg.norm(d) + 1e-12)
|
|
|
|
|
|
def rectify_candidates(vpA, vpB, L, inA, inB, shape):
|
|
"""Maps that turn the perspective lattice into an axis-aligned regular
|
|
one, so spacing becomes a 2-parameter fit instead of a projective one.
|
|
|
|
Both a projective and an affine variant are offered. Insisting on the
|
|
projective one is wrong: with a near-overhead camera the vanishing
|
|
points are far away and ill-conditioned, so the horizon estimate is
|
|
noise and the "horizon crosses the frame" guard fires on the *easiest*
|
|
inputs. Mild perspective must be the easy case.
|
|
"""
|
|
out = []
|
|
cand_hp = []
|
|
horizon = np.cross(vpA, vpB)
|
|
if abs(horizon[2]) > 1e-9:
|
|
hz = horizon / horizon[2]
|
|
Hp = np.array([[1, 0, 0], [0, 1, 0], [hz[0], hz[1], 1.0]], np.float64)
|
|
h, w = shape[:2]
|
|
corners = np.array([[0, w, w, 0], [0, 0, h, h], [1, 1, 1, 1]], np.float64)
|
|
ws = (Hp @ corners)[2]
|
|
if np.all(np.abs(ws) > 1e-6) and not (ws.min() < 0 < ws.max()):
|
|
cand_hp.append(Hp)
|
|
cand_hp.append(np.eye(3))
|
|
for Hp in cand_hp:
|
|
dA, dB = _dir_of(Hp, L[inA]), _dir_of(Hp, L[inB])
|
|
if dA is None or dB is None or abs(float(np.cross(dA, dB))) < 0.05:
|
|
continue
|
|
Ha = np.eye(3)
|
|
Ha[:2, :2] = np.linalg.inv(np.column_stack([dA, dB]))
|
|
out.append(Ha @ Hp)
|
|
return out
|
|
|
|
|
|
def fit_lattice_1d(pos, min_occupancy=0.5, top=6):
|
|
"""Fit pos ~ phase + spacing * k for unknown integers k.
|
|
|
|
Two traps, both hit for real:
|
|
|
|
* spacing -> 0 fits ANY set of positions: every value lands within
|
|
tolerance of some multiple of a tiny spacing, so maximising inlier
|
|
count collapses to a degenerate near-zero spacing. The guard is
|
|
OCCUPANCY -- the fraction of integer slots between the extreme
|
|
indices that are actually populated. A true grid fills nearly all of
|
|
them.
|
|
* fitting over every position lets ONE misdetected line (a film-strip
|
|
edge, a dotted front line) drag spacing and phase, producing a grid
|
|
visibly off by a line. So refit on inliers only.
|
|
|
|
Returns up to `top` candidates as (n_inliers, spacing, phase, occupancy),
|
|
for the caller to choose between jointly across both axes.
|
|
"""
|
|
pos = np.sort(np.asarray(pos, np.float64))
|
|
if len(pos) < 3:
|
|
return []
|
|
diffs = np.diff(pos)
|
|
diffs = diffs[diffs > 1e-9]
|
|
if len(diffs) == 0:
|
|
return []
|
|
cands = {float(np.median(diffs))}
|
|
for d in diffs:
|
|
for div in (1, 2, 3):
|
|
cands.add(d / div)
|
|
out = []
|
|
for s0 in sorted(cands):
|
|
s, phase = s0, pos[0]
|
|
if s <= 1e-9:
|
|
continue
|
|
for _ in range(3):
|
|
k = np.round((pos - phase) / s)
|
|
sol, *_ = np.linalg.lstsq(np.column_stack([np.ones(len(pos)), k]),
|
|
pos, rcond=None)
|
|
phase, s = float(sol[0]), float(sol[1])
|
|
if s <= 1e-9:
|
|
break
|
|
if s <= 1e-9:
|
|
continue
|
|
inl = np.abs(pos - (phase + np.round((pos - phase) / s) * s)) < 0.2 * s
|
|
if inl.sum() < 3:
|
|
continue
|
|
for _ in range(3): # refit on inliers only
|
|
kk = np.round((pos[inl] - phase) / s)
|
|
if len(np.unique(kk)) < 2:
|
|
break
|
|
sol, *_ = np.linalg.lstsq(np.column_stack([np.ones(int(inl.sum())), kk]),
|
|
pos[inl], rcond=None)
|
|
p2, s2 = float(sol[0]), float(sol[1])
|
|
if s2 <= 1e-9:
|
|
break
|
|
phase, s = p2, s2
|
|
nxt = np.abs(pos - (phase + np.round((pos - phase) / s) * s)) < 0.2 * s
|
|
if nxt.sum() < 3 or np.array_equal(nxt, inl):
|
|
break
|
|
inl = nxt
|
|
k = np.round((pos - phase) / s)
|
|
inl = np.abs(pos - (phase + k * s)) < 0.2 * s
|
|
if inl.sum() < 3:
|
|
continue
|
|
ks = np.unique(k[inl])
|
|
slots = ks.max() - ks.min() + 1
|
|
occ = len(ks) / slots if slots > 0 else 0.0
|
|
if occ < min_occupancy:
|
|
continue
|
|
out.append((int(inl.sum()), float(s), float(phase), float(occ)))
|
|
out.sort(key=lambda r: (-r[0], -r[1]))
|
|
keep = []
|
|
for r in out:
|
|
if all(abs(r[1] - k[1]) > 0.03 * max(r[1], k[1]) for k in keep):
|
|
keep.append(r)
|
|
if len(keep) >= top:
|
|
break
|
|
return keep
|
|
|
|
|
|
def _cluster(vals, tol):
|
|
vals = np.sort(np.asarray(vals, np.float64))
|
|
out, grp = [], [vals[0]]
|
|
for v in vals[1:]:
|
|
if v - grp[-1] <= tol:
|
|
grp.append(v)
|
|
else:
|
|
out.append(float(np.mean(grp)))
|
|
grp = [v]
|
|
out.append(float(np.mean(grp)))
|
|
return np.array(out)
|
|
|
|
|
|
def cell_steps(H, shape):
|
|
"""Pixel length of a one-index step along each lattice axis, measured at
|
|
the CENTRE of the frame -- not at index (0,0), which is usually far
|
|
off-screen and, under perspective, a wildly different scale."""
|
|
h, w = shape[:2]
|
|
c = np.linalg.inv(H) @ np.array([w / 2.0, h / 2.0, 1.0])
|
|
if abs(c[2]) < 1e-12:
|
|
return None
|
|
ci, cj = c[0] / c[2], c[1] / c[2]
|
|
q = H @ np.array([[ci, ci + 1, ci], [cj, cj, cj + 1], [1, 1, 1.0]])
|
|
if np.any(np.abs(q[2]) < 1e-12):
|
|
return None
|
|
q = q[:2] / q[2]
|
|
return (float(np.linalg.norm(q[:, 1] - q[:, 0])),
|
|
float(np.linalg.norm(q[:, 2] - q[:, 0])))
|
|
|
|
|
|
def lattice_candidates(img, L, inA, inB, top=6):
|
|
"""Candidate homographies mapping lattice index -> image pixels.
|
|
|
|
Includes BOTH axis assignments: the two line families are unordered, and
|
|
getting this wrong yields the true cell transposed.
|
|
|
|
The two axes are also chosen jointly, not independently: cells are square
|
|
on the table and the camera is near overhead, so a reconstructed cell must
|
|
come out roughly rectangular on screen. Choosing per-axis lets one axis
|
|
lock to the 1 km grid while the other locks to the 100 m subgrid, giving a
|
|
geometrically impossible 10:1 cell.
|
|
"""
|
|
vpA, vpB = _fit_vp(L, inA), _fit_vp(L, inB)
|
|
out = []
|
|
for Hr in rectify_candidates(vpA, vpB, L, inA, inB, img.shape):
|
|
posA, posB = [], []
|
|
for sub, axis, acc in ((L[inA], 1, posA), (L[inB], 0, posB)):
|
|
for (x1, y1, x2, y2) in sub:
|
|
p = Hr @ np.array([[x1, x2], [y1, y2], [1.0, 1.0]])
|
|
if np.any(np.abs(p[2]) < 1e-9):
|
|
continue
|
|
p = p[:2] / p[2]
|
|
acc.append(float(np.mean(p[axis])))
|
|
if len(posA) < 3 or len(posB) < 3:
|
|
continue
|
|
cA = fit_lattice_1d(_cluster(posA, 0.01 * max(np.ptp(posA), 1e-9)))
|
|
cB = fit_lattice_1d(_cluster(posB, 0.01 * max(np.ptp(posB), 1e-9)))
|
|
for nA, sA, pA, _oa in cA:
|
|
for nB, sB, pB, _ob in cB:
|
|
K = np.array([[sB, 0, pB], [0, sA, pA], [0, 0, 1.0]])
|
|
H = np.linalg.inv(Hr) @ K
|
|
st = cell_steps(H, img.shape)
|
|
if st is None or min(st) < 2.0 or not (0.5 <= st[0] / st[1] <= 2.0):
|
|
continue
|
|
for swap in (False, True):
|
|
HH = H @ SWAP if swap else H
|
|
out.append((HH, nA + nB, cell_steps(HH, img.shape) or st))
|
|
out.sort(key=lambda r: -r[1])
|
|
return out[:top]
|
|
|
|
|
|
# ------------------------------------------------------------- label reading
|
|
|
|
_TEMPLATES: dict = {}
|
|
|
|
|
|
def glyph_template(label: str, height: int):
|
|
"""The label as the game draws it: cream glyphs with a heavy dark
|
|
outline. The outline is what makes correlation discriminative against
|
|
aerial-photo texture, which has plenty of bright blobs but nothing
|
|
ringed in near-black."""
|
|
key = (label, height)
|
|
if key in _TEMPLATES:
|
|
return _TEMPLATES[key]
|
|
font = ImageFont.truetype(str(FONT_PATH), int(height))
|
|
pad = int(height * 0.6)
|
|
im = Image.new("L", (int(height * 5) + pad, int(height * 2) + pad), 0)
|
|
ImageDraw.Draw(im).text((pad // 2, pad // 4), label, font=font, fill=255,
|
|
stroke_width=max(1, int(height * 0.10)), stroke_fill=0)
|
|
a = np.array(im)
|
|
ys, xs = np.nonzero(a > 40)
|
|
if len(xs) == 0:
|
|
_TEMPLATES[key] = None
|
|
return None
|
|
a = a[max(0, ys.min() - 2):ys.max() + 3, max(0, xs.min() - 2):xs.max() + 3]
|
|
_TEMPLATES[key] = a
|
|
return a
|
|
|
|
|
|
def rectify_cell(img, H, i, j):
|
|
"""Warp one lattice cell to a canonical square, with a margin.
|
|
|
|
The margin matters: the lattice phase can be off by ~10% of a cell, and
|
|
an exact-bounds warp would clip a label near the cell edge -- a clipped
|
|
glyph correlates with nothing. Reading a slightly larger region tolerates
|
|
that instead of requiring the phase to be perfect.
|
|
"""
|
|
m = CELL_MARGIN
|
|
src = np.array([[i - m, j - m], [i + 1 + m, j - m],
|
|
[i + 1 + m, j + 1 + m], [i - m, j + 1 + m]], np.float64)
|
|
q = H @ np.vstack([src.T, np.ones(4)])
|
|
if np.any(np.abs(q[2]) < 1e-9):
|
|
return None
|
|
n = int(CELL_PX * (1 + 2 * m))
|
|
canon = np.array([[0, 0], [n, 0], [n, n], [0, n]], np.float32)
|
|
M = cv2.getPerspectiveTransform((q[:2] / q[2]).T.astype(np.float32), canon)
|
|
return cv2.warpPerspective(img, M, (n, n), flags=cv2.INTER_LINEAR)
|
|
|
|
|
|
def read_cell_label(cell_gray, glyph_fracs=(0.10, 0.13, 0.17)):
|
|
"""Which label best explains the pixels where a label must be?
|
|
|
|
No detection: the grid fixes the label's position and size, so this
|
|
correlates every candidate label there and takes the best. Sliding the
|
|
template over a slightly larger crop absorbs the residual phase error.
|
|
All four cell orientations are tried, because the lattice axes have
|
|
arbitrary direction and the label may land in any corner, mirrored.
|
|
|
|
-> (label, score) with score in [-1, 1]; measured, correct reads score
|
|
0.73-0.87 and wrong ones 0.40-0.56.
|
|
"""
|
|
n = cell_gray.shape[0]
|
|
best = (None, -1.0)
|
|
x0 = int(max(0, (CELL_MARGIN - 0.03) * CELL_PX))
|
|
x1 = int(min(n, (CELL_MARGIN + 0.50) * CELL_PX))
|
|
y0 = int(max(0, (CELL_MARGIN - 0.05) * CELL_PX))
|
|
y1 = int(min(n, (CELL_MARGIN + 0.30) * CELL_PX))
|
|
for fx, fy in FLIPS:
|
|
v = cell_gray
|
|
if fx:
|
|
v = v[:, ::-1]
|
|
if fy:
|
|
v = v[::-1, :]
|
|
patch = np.ascontiguousarray(v[y0:y1, x0:x1])
|
|
if patch.shape[0] < 12 or patch.shape[1] < 12:
|
|
continue
|
|
for gf in glyph_fracs:
|
|
h = max(8, int(gf * CELL_PX))
|
|
for lab in LABELS:
|
|
t = glyph_template(lab, h)
|
|
if t is None or t.shape[0] >= patch.shape[0] or t.shape[1] >= patch.shape[1]:
|
|
continue
|
|
sc = float(cv2.matchTemplate(patch, t, cv2.TM_CCOEFF_NORMED).max())
|
|
if sc > best[1]:
|
|
best = (lab, sc)
|
|
return best
|
|
|
|
|
|
def visible_cells(H, shape, limit=6):
|
|
"""Lattice cells whose centre is on screen, nearest the frame centre
|
|
first (least perspective distortion, so the easiest to read)."""
|
|
h, w = shape[:2]
|
|
q = np.linalg.inv(H) @ np.array([[0, w, w, 0], [0, 0, h, h], [1, 1, 1, 1]], np.float64)
|
|
ok = np.abs(q[2]) > 1e-9
|
|
if ok.sum() < 3:
|
|
return []
|
|
ij = q[:2, ok] / q[2, ok]
|
|
cells = []
|
|
for i in range(int(np.floor(ij[0].min())), int(np.ceil(ij[0].max()))):
|
|
for j in range(int(np.floor(ij[1].min())), int(np.ceil(ij[1].max()))):
|
|
c = H @ np.array([i + 0.5, j + 0.5, 1.0])
|
|
if abs(c[2]) < 1e-9:
|
|
continue
|
|
x, y = c[0] / c[2], c[1] / c[2]
|
|
if 0 <= x < w and 0 <= y < h:
|
|
cells.append((float(np.hypot(x - w / 2, y - h / 2)), i, j))
|
|
cells.sort()
|
|
return [(i, j) for _d, i, j in cells[:limit]]
|
|
|
|
|
|
# ------------------------------------------------------------------- solving
|
|
|
|
class GridSolution:
|
|
"""A solved grid: lattice homography plus the discrete mapping from
|
|
lattice index to game cell."""
|
|
|
|
def __init__(self, H, si, sj, du, dv, votes, reads, steps):
|
|
self.H, self.si, self.sj = H, si, sj
|
|
self.du, self.dv = du, dv
|
|
self.votes, self.reads, self.steps = votes, reads, steps
|
|
|
|
def lattice_to_grid(self):
|
|
"""Affine 3x3 taking lattice coords (u, v) to continuous game grid
|
|
coords (col, row), col in [0, 20] and row in [1, 11].
|
|
|
|
The +1 for a negative sign is not cosmetic. With si = -1, increasing u
|
|
means decreasing col, so the cell that lattice index i opens at u = i
|
|
is entered from its RIGHT edge, and si*u + du alone runs from col+1
|
|
down to col across it -- floor() would name the neighbour for the whole
|
|
cell. Offsetting by 1 makes the fraction always grow in the direction
|
|
col/row grow, which is also what the app's Coord means by its sub-cell
|
|
x/y (see models.Coord.as_fraction), so both signs agree with it.
|
|
"""
|
|
return np.array([[self.si, 0.0, self.du + (0 if self.si > 0 else 1)],
|
|
[0.0, self.sj, self.dv + (0 if self.sj > 0 else 1)],
|
|
[0.0, 0.0, 1.0]])
|
|
|
|
def grid_of(self, x, y):
|
|
"""-> continuous (col, row) for a pixel in working-resolution coords."""
|
|
q = self.lattice_to_grid() @ np.linalg.inv(self.H) @ np.array([x, y, 1.0])
|
|
if abs(q[2]) < 1e-12:
|
|
return None
|
|
return float(q[0] / q[2]), float(q[1] / q[2])
|
|
|
|
def cell_of(self, x, y):
|
|
"""-> ("J8", sub_x, sub_y) for a pixel in working-resolution coords."""
|
|
g = self.grid_of(x, y)
|
|
if g is None:
|
|
return None
|
|
colf, rowf = g
|
|
col, row = int(np.floor(colf)), int(np.floor(rowf))
|
|
if not (0 <= col < COLS and 1 <= row <= ROWS):
|
|
return None
|
|
return (f"{LARGE_X[col]}{row}",
|
|
int(np.clip((colf - col) * 10, 0, 9)),
|
|
int(np.clip((rowf - row) * 10, 0, 9)))
|
|
|
|
|
|
def format_coord(cell) -> str:
|
|
"""("K8", 0, 3) -> "K8 0:3", matching how the game writes coordinates."""
|
|
if cell is None:
|
|
return "?"
|
|
return f"{cell[0]} {cell[1]}:{cell[2]}"
|
|
|
|
|
|
def solve(img):
|
|
"""-> (GridSolution, None) or (None, reason)."""
|
|
L = segments(ridge(img), min_len=0.04 * img.shape[1])
|
|
inA, inB = two_families(L)
|
|
if inA.sum() < 3 or inB.sum() < 3:
|
|
return None, "too few grid line families"
|
|
cands = lattice_candidates(img, L, inA, inB)
|
|
if not cands:
|
|
return None, "no plausible lattice"
|
|
best = None
|
|
for H, _ninl, steps in cands:
|
|
cells = visible_cells(H, img.shape)
|
|
reads = []
|
|
for (i, j) in cells:
|
|
cell = rectify_cell(img, H, i, j)
|
|
if cell is None:
|
|
continue
|
|
lab, sc = read_cell_label(cv2.cvtColor(cell, cv2.COLOR_BGR2GRAY))
|
|
if lab and sc >= LABEL_ACCEPT:
|
|
reads.append((i, j, lab, sc))
|
|
if len(reads) < MIN_LABEL_VOTES:
|
|
continue
|
|
for si in (1, -1):
|
|
for sj in (1, -1):
|
|
votes = {}
|
|
for i, j, lab, sc in reads:
|
|
key = (LARGE_X.index(lab[0]) - si * i, int(lab[1:]) - sj * j)
|
|
v = votes.setdefault(key, [0, 0.0])
|
|
v[0] += 1
|
|
v[1] += sc
|
|
for (du, dv), (cnt, tot) in votes.items():
|
|
if cnt < MIN_LABEL_VOTES:
|
|
continue
|
|
# extent prior: every visible cell must be a real map cell
|
|
inrange = all(0 <= si * i + du < COLS and 1 <= sj * j + dv <= ROWS
|
|
for (i, j) in cells)
|
|
score = tot + cnt + (1.5 if inrange else -1.5)
|
|
if best is None or score > best[0]:
|
|
best = (score, H, si, sj, du, dv, cnt, len(reads), steps)
|
|
if best is None:
|
|
return None, (f"no confident label read "
|
|
f"(need {MIN_LABEL_VOTES} at >={LABEL_ACCEPT})")
|
|
_s, H, si, sj, du, dv, cnt, nreads, steps = best
|
|
return GridSolution(H, si, sj, du, dv, cnt, nreads, steps), None
|
|
|
|
|
|
def solution_from_correspondences(pairs):
|
|
"""Build a solution from explicit grid<->pixel correspondences.
|
|
|
|
`pairs` is [((col, row), (x, y)), ...] with at least 4 entries, where
|
|
(col, row) are CONTINUOUS grid coordinates: col 0..20 increasing with the
|
|
letters, row 1..11 as the game numbers them. This is the manual override
|
|
path -- four dragged cell corners plus that cell's label fully determine
|
|
the homography, so it works even when line detection or label reading
|
|
fail completely.
|
|
|
|
Four correspondences is the minimum: a homography has 8 degrees of
|
|
freedom and each point contributes 2 equations. Three points would only
|
|
fix an affine map, and under real perspective a square's image is a
|
|
general quadrilateral, so the fourth corner is genuinely not implied by
|
|
the other three.
|
|
"""
|
|
if len(pairs) < 4:
|
|
raise ValueError("a homography needs at least 4 correspondences")
|
|
src = np.array([[p[0][0], p[0][1]] for p in pairs], np.float64)
|
|
dst = np.array([[p[1][0], p[1][1]] for p in pairs], np.float64)
|
|
H, _ = cv2.findHomography(src.reshape(-1, 1, 2), dst.reshape(-1, 1, 2), 0)
|
|
if H is None:
|
|
raise ValueError("degenerate correspondences")
|
|
# identity discrete mapping: the grid coords were given directly
|
|
return GridSolution(H, 1, 1, 0, 0, votes=len(pairs), reads=len(pairs),
|
|
steps=cell_steps(H, (1, 1)) or (1.0, 1.0))
|
|
|
|
|
|
def centre_cell_quad(sol, shape):
|
|
"""The centre-most fully-visible cell, as handles for manual correction.
|
|
|
|
-> (label, [(x, y) x4], [(col, row) x4]) with the two lists in matching
|
|
order, so a UI can seed four draggable handles from the automatic
|
|
solution and refit through solution_from_correspondences() as they move.
|
|
"""
|
|
h, w = shape[:2]
|
|
cells = visible_cells(sol.H, shape, limit=1)
|
|
if not cells:
|
|
return None
|
|
i, j = cells[0]
|
|
corners_ij = [(i, j), (i + 1, j), (i + 1, j + 1), (i, j + 1)]
|
|
q = sol.H @ np.array([[c[0] for c in corners_ij],
|
|
[c[1] for c in corners_ij], [1, 1, 1, 1.0]])
|
|
if np.any(np.abs(q[2]) < 1e-9):
|
|
return None
|
|
px = [(float(x), float(y)) for x, y in (q[:2] / q[2]).T]
|
|
g = sol.lattice_to_grid() @ np.array([[c[0] for c in corners_ij],
|
|
[c[1] for c in corners_ij], [1, 1, 1, 1.0]])
|
|
grid = [(float(a), float(b)) for a, b in (g[:2] / g[2]).T]
|
|
centre = sol.cell_of(*((np.array(px[0]) + np.array(px[2])) / 2))
|
|
return (centre[0] if centre else None), px, grid
|
|
|
|
|
|
MAP_KM_W, MAP_KM_H = 20.0, 10.0
|
|
|
|
|
|
def warp_to_map(img, sol, px_per_km=100):
|
|
"""Rectify a screenshot into map space, ready to composite under the app's
|
|
own grid.
|
|
|
|
Returns (BGRA array, px_per_km). Only the region the screenshot actually
|
|
covers is opaque; everything else is transparent, so a partial view of the
|
|
table does not blank out the rest of the map.
|
|
|
|
Warping once into map space -- rather than transforming while drawing --
|
|
keeps the renderer simple: Cairo has no projective transform, but once the
|
|
image is in map space a plain scale and translate places it.
|
|
"""
|
|
out_w, out_h = int(MAP_KM_W * px_per_km), int(MAP_KM_H * px_per_km)
|
|
# Map space is pixels over the whole board: x = col * px_per_km rightward,
|
|
# y measured DOWN while row counts UP, so row 1 (the game's bottom row)
|
|
# lands at the bottom edge. Row is 1-based here and 0-based in the app's
|
|
# Coord.as_fraction, hence the extra +px_per_km.
|
|
grid_to_map = np.array([[px_per_km, 0.0, 0.0],
|
|
[0.0, -px_per_km, out_h + px_per_km],
|
|
[0.0, 0.0, 1.0]])
|
|
# sol.H alone only reaches LATTICE coords; the discrete mapping (si, sj,
|
|
# du, dv) is what pins those to named cells, and leaving it out put the
|
|
# screenshot in the wrong place for every automatically solved grid.
|
|
M = grid_to_map @ sol.lattice_to_grid() @ np.linalg.inv(sol.H)
|
|
bgra = cv2.cvtColor(img, cv2.COLOR_BGR2BGRA)
|
|
bgra[:, :, 3] = 255
|
|
return cv2.warpPerspective(bgra, M, (out_w, out_h), flags=cv2.INTER_LINEAR,
|
|
borderMode=cv2.BORDER_CONSTANT,
|
|
borderValue=(0, 0, 0, 0)), px_per_km
|
|
|
|
|
|
RETRY_WORK_W = 2400
|
|
|
|
|
|
GATE_MIN_FAMILY = 8
|
|
|
|
|
|
def looks_like_map(img) -> bool:
|
|
"""Cheap routing test: is this a map-table screenshot or writer text?
|
|
|
|
Purely a latency optimisation, NOT a correctness gate -- `solve()` is the
|
|
real decision, and measured over 122 writer screenshots it accepts zero of
|
|
them. This just avoids paying solve()'s ~10-20s on text pastes.
|
|
|
|
The discriminating feature is the number of lines in the SMALLER line
|
|
family: measured, map screenshots have >=9 (median 15) while writer
|
|
screenshots that reach this point have a median of 5 (lower quartile 3).
|
|
Angular separation and implied cell size do not separate the two at all.
|
|
|
|
Deliberately looser than solve(): a false positive here only costs time,
|
|
while a false negative would silently route a map to the text pipeline.
|
|
"""
|
|
L = segments(ridge(img), min_len=0.04 * img.shape[1])
|
|
inA, inB = two_families(L)
|
|
if min(int(inA.sum()), int(inB.sum())) < GATE_MIN_FAMILY:
|
|
return False
|
|
return bool(lattice_candidates(img, L, inA, inB))
|
|
|
|
|
|
def solve_path(path):
|
|
"""Solve a screenshot on disk, retrying at higher working resolution.
|
|
|
|
A very wide screenshot downscaled to WORK_W can leave cells so small that
|
|
the label is only a handful of pixels across, which no filtering recovers
|
|
-- measured, one fixture's best label score goes 0.54 -> 0.65 (past the
|
|
accept threshold) purely from working at 2400px instead of 1500px. The
|
|
retry only pays that cost when the first pass actually fails.
|
|
|
|
-> (GridSolution, image_used, None) or (None, image_used, reason)
|
|
"""
|
|
img = load(path)
|
|
sol, err = solve(img)
|
|
if sol is not None:
|
|
return sol, img, None
|
|
raw = cv2.imread(str(path), cv2.IMREAD_COLOR)
|
|
if raw is not None and raw.shape[1] > 1.4 * WORK_W:
|
|
big = downscale(raw, RETRY_WORK_W)
|
|
sol2, err2 = solve(big)
|
|
if sol2 is not None:
|
|
return sol2, big, None
|
|
err = err2 or err
|
|
return None, img, err
|
|
|
|
|
|
# ------------------------------------------------------------------- markers
|
|
|
|
# Grey-world illuminant normalisation was tried here and removed. It did cut
|
|
# one fixture's false positives (33 -> 5) by neutralising red stage lighting,
|
|
# but it made the worst case worse: cancelling the red cast also restores the
|
|
# cyan front-line ribbon to full saturation, so the ribbon then fires the
|
|
# friendly-marker mask (43 -> 50 spurious markers). The false positives are
|
|
# overlay geometry, not lighting, so they need a shape test, not a colour fix.
|
|
|
|
|
|
# ---- unit type classification against the game's own marker icons --------
|
|
|
|
_ICON_BANK: dict = {}
|
|
ICON_SIZE = 64
|
|
DIAMOND_IOU = 0.64 # blob-vs-ideal-diamond overlap needed to be a marker.
|
|
# Swept against verified counts: 0.64 keeps every shot
|
|
# confirmed correct by hand (5/2/2/3 markers) while cutting
|
|
# ribbon+hatching false positives from 43 to 2 on the worst
|
|
# fixture. Loosening to 0.50 regains one real marker on one
|
|
# shot but quadruples the false positives.
|
|
SYMBOL_KEEP = 0.52 # central fraction of the marker that carries the symbol
|
|
|
|
|
|
def _icon_dir(side):
|
|
base = Path(__file__).resolve().parents[2] / "assets" / "icons" / "targets"
|
|
return base / ("friendly" if side == "friendly" else "enemy")
|
|
|
|
|
|
def icon_bank(side):
|
|
"""Every marker icon as a normalised grayscale patch.
|
|
|
|
The shipped icons are the complete marker -- coloured diamond plus the
|
|
black inner symbol -- which is exactly what is drawn on the table, so a
|
|
detected marker can be matched against them directly. They differ ONLY in
|
|
the inner symbol, so the comparison is effectively on that symbol.
|
|
"""
|
|
if side in _ICON_BANK:
|
|
return _ICON_BANK[side]
|
|
entries = []
|
|
d = _icon_dir(side)
|
|
for f in sorted(d.glob("*.png")):
|
|
raw = cv2.imread(str(f), cv2.IMREAD_UNCHANGED)
|
|
if raw is None:
|
|
continue
|
|
if raw.shape[2] == 4:
|
|
a = raw[:, :, 3:4].astype(np.float32) / 255.0
|
|
rgb = raw[:, :, :3].astype(np.float32)
|
|
raw = (rgb * a + 128.0 * (1 - a)).astype(np.uint8)
|
|
g = cv2.cvtColor(raw, cv2.COLOR_BGR2GRAY)
|
|
g = cv2.resize(g, (ICON_SIZE, ICON_SIZE), interpolation=cv2.INTER_AREA)
|
|
name = f.stem.split("_", 1)[1] if "_" in f.stem else f.stem
|
|
entries.append((name, _zscore(_inner(g))))
|
|
_ICON_BANK[side] = entries
|
|
return entries
|
|
|
|
|
|
def _inner(patch):
|
|
"""The central part of a marker, where the only discriminative content is.
|
|
|
|
Every icon is the SAME diamond and differs only in the small black symbol
|
|
inside it, so correlating whole markers lets the identical diamond edges
|
|
dominate the score and swamp the signal -- which is why classification
|
|
returned near-tied scores (margins under 0.05) and therefore almost always
|
|
None.
|
|
"""
|
|
n = patch.shape[0]
|
|
k = int(n * (1 - SYMBOL_KEEP) / 2)
|
|
return patch[k:n - k, k:n - k]
|
|
|
|
|
|
def _zscore(patch):
|
|
p = patch.astype(np.float32)
|
|
p -= p.mean()
|
|
sd = float(p.std())
|
|
return p / sd if sd > 1e-6 else p
|
|
|
|
|
|
def classify_marker(img, box, side, pad=0.22, min_margin=0.08, min_score=0.55):
|
|
"""Which unit icon is this marker? -> (name, score, margin).
|
|
|
|
`name` is None when the best match does not beat the runner-up by
|
|
`min_margin`; the markers are small on screen and several icons differ
|
|
only in fine detail, so an unconfident answer must stay unknown rather
|
|
than become a wrong unit type.
|
|
"""
|
|
x, y, w, h = box
|
|
m = int(pad * max(w, h))
|
|
x0, y0 = max(0, x - m), max(0, y - m)
|
|
x1, y1 = min(img.shape[1], x + w + m), min(img.shape[0], y + h + m)
|
|
crop = img[y0:y1, x0:x1]
|
|
if crop.size == 0 or min(crop.shape[:2]) < 8:
|
|
return None, 0.0, 0.0
|
|
g = cv2.resize(cv2.cvtColor(crop, cv2.COLOR_BGR2GRAY),
|
|
(ICON_SIZE, ICON_SIZE), interpolation=cv2.INTER_AREA)
|
|
q = _zscore(_inner(g))
|
|
scores = [(float((q * t).mean()), name) for name, t in icon_bank(side)]
|
|
if not scores:
|
|
return None, 0.0, 0.0
|
|
scores.sort(reverse=True)
|
|
best, second = scores[0], (scores[1] if len(scores) > 1 else (0.0, None))
|
|
margin = best[0] - second[0]
|
|
# An absolute floor as well as a margin. With only the margin, matching
|
|
# collapsed onto one class ("Ship") for nearly every marker at scores of
|
|
# 0.29-0.46 -- confidently wrong, which is worse than admitting ignorance,
|
|
# because a wrong unit type looks like real intel. Markers are only ~30px
|
|
# across at the working resolution; cropping the symbol from the NATIVE
|
|
# resolution screenshot is the fix, not a lower threshold.
|
|
ok = margin >= min_margin and best[0] >= min_score
|
|
return (best[1] if ok else None), best[0], margin
|
|
|
|
|
|
def marker_masks(img):
|
|
"""Hostile (pink/red) and friendly (cyan) marker colours."""
|
|
hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)
|
|
h, s, v = hsv[:, :, 0], hsv[:, :, 1], hsv[:, :, 2]
|
|
hostile = ((h <= 10) | (h >= 168)) & (s > 55) & (v > 95)
|
|
friendly = (h >= 82) & (h <= 105) & (s > 55) & (v > 110)
|
|
k = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (3, 3))
|
|
return [cv2.morphologyEx(m.astype(np.uint8) * 255, cv2.MORPH_CLOSE, k)
|
|
for m in (hostile, friendly)]
|
|
|
|
|
|
def diamonds(mask, cell_px):
|
|
"""Marker-sized, marker-shaped blobs.
|
|
|
|
The markers scale with the map, so a solved grid tells us how big one
|
|
must be (~0.14 of a cell). A diamond also fills about half its bounding
|
|
box, which rejects the long thin territory hatching and front-line
|
|
ribbons that share the markers' colours.
|
|
"""
|
|
want = 0.14 * cell_px
|
|
lo, hi = 0.55 * want, 1.9 * want
|
|
n, lab, stats, cent = cv2.connectedComponentsWithStats(mask, 8)
|
|
out = []
|
|
for i in range(1, n):
|
|
x, y, w, h, a = stats[i]
|
|
if not (lo <= max(w, h) <= hi) or min(w, h) < 0.4 * lo:
|
|
continue
|
|
if not (0.55 <= w / h <= 1.8):
|
|
continue
|
|
if not (0.30 <= a / float(w * h) <= 0.85):
|
|
continue
|
|
# Actually test for a DIAMOND. A bounding-box fill ratio near 0.5 is
|
|
# not enough: a chunk of the territory hatching or of a front-line
|
|
# ribbon hits the same ratio and the same colour, which is where the
|
|
# tens of spurious markers came from. Compare the blob against an
|
|
# ideal diamond inscribed in its own bounding box.
|
|
blob = (lab[y:y + h, x:x + w] == i)
|
|
ideal = np.zeros((h, w), np.uint8)
|
|
cv2.fillConvexPoly(ideal, np.array(
|
|
[[w // 2, 0], [w - 1, h // 2], [w // 2, h - 1], [0, h // 2]], np.int32), 1)
|
|
ideal = ideal.astype(bool)
|
|
union = int(np.logical_or(blob, ideal).sum())
|
|
if union == 0:
|
|
continue
|
|
if int(np.logical_and(blob, ideal).sum()) / union < DIAMOND_IOU:
|
|
continue
|
|
out.append((float(cent[i][0]), float(cent[i][1]), (int(x), int(y), int(w), int(h))))
|
|
return out
|
|
|
|
|
|
def find_markers(img, sol):
|
|
"""-> list of dicts: side, unit, label, sub_x, sub_y, coord, centre, box."""
|
|
cell = max(sol.steps)
|
|
found = []
|
|
for side, mask in zip(("hostile", "friendly"), marker_masks(img)):
|
|
for (cx, cy, box) in diamonds(mask, cell):
|
|
c = sol.cell_of(cx, cy)
|
|
if c is None:
|
|
continue
|
|
unit, score, margin = classify_marker(img, box, side)
|
|
found.append(dict(side=side, unit=unit, unit_score=score,
|
|
unit_margin=margin, label=c[0], sub_x=c[1],
|
|
sub_y=c[2], coord=format_coord(c),
|
|
centre=(cx, cy), box=box))
|
|
return found
|