Sweep em-dashes out of code, docstrings, and requirements.txt
Same style fix applied to the README earlier, extended everywhere: replaced " -- " with commas/colons/periods (picking whichever reads right per occurrence, splitting into two sentences where the clauses were independent), fixed a few user-facing strings along the way (entity list rows, placement/strike toasts, ambiguous-candidate tag, shell picker button label). Left three intentional non-prose uses alone: the "unassigned" dash glyph in firing_panel.py (and its docstring diagram), and ocr.py's dash-variant regex character class, which needs to literally match em/en-dashes in OCR'd text. Also caught and fixed a stale models.py docstring claiming "no solver yet" (solver.py has existed for a while) while touching that paragraph anyway, and a formatting artifact in coord_dialog.py's docstring left by the sed pass (misaligned comma from a since-removed alignment gap). Verified: py_compile across all files, the 9-screenshot OCR regression sweep, and a GTK smoke test exercising the edited toast/placement code paths. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
+20
-20
@@ -1,12 +1,12 @@
|
||||
"""Text OCR pipeline: screenshot -> cleaned-up text -> parsed board info.
|
||||
|
||||
Only the text sub-pipeline is implemented. There's no image/icon
|
||||
recognition sub-pipeline yet (spotting markers, ship icons, etc.) — that's
|
||||
recognition sub-pipeline yet (spotting markers, ship icons, etc.), that's
|
||||
a separate future pipeline, out of scope here.
|
||||
|
||||
Preprocessing matters more than the regexes: the typewriter photo has an
|
||||
uneven vignette (in-game light falloff) that a single global threshold
|
||||
can't handle — it either loses faint corners or blobs-out dark ones. We
|
||||
can't handle, it either loses faint corners or blobs-out dark ones. We
|
||||
flatten that by dividing by a heavily blurred copy of itself (crude local
|
||||
background normalization) before thresholding, which recovers text in the
|
||||
darkened areas reliably.
|
||||
@@ -94,13 +94,13 @@ _SEP = r"\s*[-–—]\s*"
|
||||
_COORD_RE = re.compile(
|
||||
rf"([A-T])\s*({_DIGIT_CLASS}{{1,2}})\s+({_DIGIT_CLASS})\s*[:;.,]\s*({_DIGIT_CLASS})"
|
||||
)
|
||||
# No literal '#' required — it's just as OCR-corruptible as anything else
|
||||
# No literal '#' required, it's just as OCR-corruptible as anything else
|
||||
# (missing entirely, or misread as e.g. 'H'). We instead anchor to *where*
|
||||
# the fuzzy keyword match ended and take the first run of digit-shaped
|
||||
# characters after that, skipping over whatever separator survived.
|
||||
_ID_RE = re.compile(rf"({_DIGIT_CLASS}+)")
|
||||
|
||||
# Not seen in a real screenshot yet, so no extraction for these — add a
|
||||
# Not seen in a real screenshot yet, so no extraction for these, add a
|
||||
# keyword + extractor here (plus a field below and a case in parse_text)
|
||||
# once we know the format:
|
||||
# - reference points
|
||||
@@ -114,7 +114,7 @@ def _fix_digits(s: str) -> str:
|
||||
def _fix_id_digits(raw: str) -> str:
|
||||
"""Like _fix_digits, but for id-length runs specifically: also
|
||||
collapses a 2-character run where one char is a genuine digit and the
|
||||
other a look-alike letter that maps to the *same* digit — e.g. '1l' or
|
||||
other a look-alike letter that maps to the *same* digit, e.g. '1l' or
|
||||
'S5' both fix to '11'/'55', but a real 2-digit id wouldn't plausibly
|
||||
render as one numeral plus one letter of the identical value; that
|
||||
shape is the signature of OCR ghosting a single thin glyph twice
|
||||
@@ -132,7 +132,7 @@ def _fix_id_digits(raw: str) -> str:
|
||||
|
||||
# A target can also be spotted with an absolute grid ref directly
|
||||
# ("Target#10 Spotted. Grid Q3 9:0") instead of/alongside bearing/distance
|
||||
# clues — same coordinate shape as _COORD_RE, just anchored after "Grid".
|
||||
# clues, same coordinate shape as _COORD_RE, just anchored after "Grid".
|
||||
_GRID_COORD_RE = re.compile(
|
||||
rf"Grid\s+([A-T])\s*({_DIGIT_CLASS}{{1,2}})\s+({_DIGIT_CLASS})\s*[:;.,]\s*({_DIGIT_CLASS})",
|
||||
re.IGNORECASE,
|
||||
@@ -189,17 +189,17 @@ def _extract_leading_id(remainder: str) -> int | None:
|
||||
#
|
||||
# Each named entity is a block of one or more clue lines, terminated by a
|
||||
# blank line or a lone '.'. Degree signs, colons, and the 'km' unit are all
|
||||
# treated as optional/lossy — OCR drops them unpredictably.
|
||||
# treated as optional/lossy, OCR drops them unpredictably.
|
||||
|
||||
_RP_HEADER_RE = re.compile(r"Reference\s+Point\s+([A-Za-z][\w-]*)\s*:?", re.IGNORECASE)
|
||||
# '#' isn't required literally — same reasoning as the spotter-id fix: OCR
|
||||
# '#' isn't required literally, same reasoning as the spotter-id fix: OCR
|
||||
# drops it or renders it as noise (seen: '€'). Up to 2 junk characters
|
||||
# between the type word and its digits is enough slack without risking a
|
||||
# false match elsewhere.
|
||||
# ^ the junk-class run is REQUIRED (1-2 chars, not 0-2): the digit class
|
||||
# deliberately overlaps the alphabet (g/s/i/l/o/... look like digits), so
|
||||
# with a 0-width separator allowed, "Bearing" backtracks into itself —
|
||||
# word="Bearin", "digit"=its own trailing 'g' — and falsely matches as a
|
||||
# with a 0-width separator allowed, "Bearing" backtracks into itself,
|
||||
# word="Bearin", "digit"=its own trailing 'g', and falsely matches as a
|
||||
# header. Requiring real punctuation between word and digits (true of
|
||||
# every observed header: '#', a misread substitute, ...) rules that out,
|
||||
# and also stops a bare "Word 094" clue line (space only) from matching.
|
||||
@@ -209,13 +209,13 @@ _LEADING_NOISE_RE = re.compile(r"^[^A-Za-z]{1,3}(?=[A-Za-z])")
|
||||
|
||||
# Reference capture is (\S+), not (.+): references are always a single
|
||||
# token with no spaces, and being non-greedy this way is what lets
|
||||
# finditer() find more than one clue per line/block — "Bearing 118 from
|
||||
# finditer() find more than one clue per line/block, "Bearing 118 from
|
||||
# Spotter#2 & Bearing 125 from Spotter#1" needs two separate matches, and
|
||||
# a greedy (.+) would let the first one swallow the rest of the string.
|
||||
#
|
||||
# _GAP sits right before "from": plain whitespace normally, but an OCR
|
||||
# line-wrap can drop a stray junk token right at the break ("Bearing 125°"
|
||||
# / "P; from Spotter#1") — and that junk can itself contain a letter (the
|
||||
# / "P; from Spotter#1"), and that junk can itself contain a letter (the
|
||||
# 'P' above), so this isn't just non-alnum noise like the header-bullet
|
||||
# case; tolerate any single short token, not just punctuation.
|
||||
_GAP = r"[\s]*(?:\S{1,3}\s*)?"
|
||||
@@ -230,7 +230,7 @@ _CLUE_DISTANCE_RE = re.compile(rf"Distance\s*([\d.]+)\s*k?m?{_GAP}from\s+(\S+)",
|
||||
|
||||
_TYPE_BY_SHORT = {t.short: t for t in TargetType}
|
||||
# The game's typewriter has used "AmmoCache" for what's now modeled as
|
||||
# SupplyCache — treat it as the same type rather than dropping the target.
|
||||
# SupplyCache, treat it as the same type rather than dropping the target.
|
||||
_TYPE_WORD_ALIASES = {"AmmoCache": "SupplyCache"}
|
||||
|
||||
|
||||
@@ -239,10 +239,10 @@ _REF_NAMED_RE = re.compile(rf"^([A-Za-z]+)[^A-Za-z0-9\s]{{1,2}}({_DIGIT_CLASS}+)
|
||||
|
||||
def _clean_reference(raw: str) -> str:
|
||||
"""Leading name-shaped token, dropping trailing OCR noise (stray dots,
|
||||
double spaces, etc.) — reference names never contain spaces. Named
|
||||
double spaces, etc.), reference names never contain spaces. Named
|
||||
references ('Spotter#1', 'AmmoCache#2') get their digit part fixed up
|
||||
and their separator normalized to '#'; plain word references ('Alpha')
|
||||
are left untouched — don't run digit-fixing over them or real letters
|
||||
are left untouched, don't run digit-fixing over them or real letters
|
||||
like the 'l' in 'Alpha' get corrupted into '1'."""
|
||||
token = re.match(r"\S+", raw.strip())
|
||||
token = token.group(0) if token else raw.strip()
|
||||
@@ -271,7 +271,7 @@ _CLUE_PATTERNS = (
|
||||
|
||||
|
||||
def _parse_all_clues(text: str) -> list[Clue]:
|
||||
"""Every Bearing/Distance clue found anywhere in `text` — a block can
|
||||
"""Every Bearing/Distance clue found anywhere in `text`, a block can
|
||||
have several (one per clue line, or more than one on a single line
|
||||
joined with '&')."""
|
||||
clues: list[Clue] = []
|
||||
@@ -292,7 +292,7 @@ def _parse_all_clues(text: str) -> list[Clue]:
|
||||
|
||||
|
||||
def parse_clues_from_text(text: str) -> list[Clue]:
|
||||
"""Parse every Bearing/Distance clue found in free-form text — used
|
||||
"""Parse every Bearing/Distance clue found in free-form text, used
|
||||
for manually-typed descriptions in the coord dialog, sharing the
|
||||
exact same clue grammar as the OCR'd intel blocks."""
|
||||
return _parse_all_clues(text)
|
||||
@@ -318,7 +318,7 @@ def parse_intel_blocks(text: str) -> list[dict]:
|
||||
with neither a clue nor a grid coord are dropped (nothing to store).
|
||||
|
||||
Clues are extracted once per block, from the whole joined block text,
|
||||
at flush time — not accumulated line-by-line while scanning. That's
|
||||
at flush time, not accumulated line-by-line while scanning. That's
|
||||
what lets a clue split across an OCR line-wrap ("...Bearing 125°" /
|
||||
"from Spotter#1" on separate lines) or two clues on one line
|
||||
("Bearing X from A & Bearing Y from B") both resolve correctly. A
|
||||
@@ -352,7 +352,7 @@ def parse_intel_blocks(text: str) -> list[dict]:
|
||||
# defeat the column-0-anchored header regexes below. Try the line
|
||||
# as-is first, and only if that fails, retry with its first
|
||||
# whitespace-delimited token stripped (covers symbol junk *and*
|
||||
# a misread bullet that happened to OCR as a stray letter) — but
|
||||
# a misread bullet that happened to OCR as a stray letter), but
|
||||
# only when that leading token is bullet-length (<=3 chars), else
|
||||
# a genuine wrapped clue continuation like "from Spotter#2" gets
|
||||
# its "from" stripped and "Spotter#2" misread as a new header.
|
||||
@@ -390,7 +390,7 @@ def parse_intel_blocks(text: str) -> list[dict]:
|
||||
# Destruction reports are standalone one-liners, not tied to a block:
|
||||
# "SupplyCache#2 Destroyed. Additional Requisition Granted."
|
||||
# "Direct Hit! HostileTank#3 Destroyed."
|
||||
# Just need "<Type>#<id>" immediately followed by "Destroyed" — the
|
||||
# Just need "<Type>#<id>" immediately followed by "Destroyed", the
|
||||
# "Direct Hit!" prefix (or its absence) doesn't matter, search() finds
|
||||
# the name+Destroyed pair anywhere in the line either way.
|
||||
_DESTROYED_RE = re.compile(
|
||||
|
||||
Reference in New Issue
Block a user