# Follow-up plan: bounded country interventions before a visible cue

Recorded before follow-up model inference. Preserve all prior-study files. Reuse its original Qwen3-1.7B bf16 checkpoint, 466-prompt lens, environment and verified MPS setup. No explanation runs, hardware comparison, model search or lens fitting.

## Hypotheses and evidence boundary

H1: a bounded country-coordinate intervention at block 9/13/17 and one original prompt position produces a target-currency likelihood shift and sometimes a clean target-code generation before any country/nationality/other substantive cue. H2: the J transformation supplies a paired advantage over the same coordinate construction using ordinary final-normalization/unembedding geometry. H3 (natural use of this intermediate by the unmodified model) is not tested: neither H1 nor H2 proves it. Null effects are retained without grid expansion.

## Task and provenance

Eight development cities, 32 evaluation cities (two per source country). All cities are fresh relative to the prior study. Development and evaluation source/target country groups are disjoint, but the country groups are reused from the prior experiment: this is city holdout evidence. A fixed cyclic mapping gives each source country a different target currency. Task labels use current ISO 4217 codes, directly checked against SIX's downloaded list. Location records are checked against UN/LOCODE 2024-2 via a pinned public mirror because the UN download endpoints return 403; official indexed UN pages provide corroboration. All source records and access failures are retained. Conventional English city aliases map explicitly to directory entries. Avoid disputed locations and globally prominent ambiguous names such as Salvador/Merida. Spelling normalization never substitutes an unrelated location.

Ordinary unconstrained generation. Three declared prompt candidates use the pinned tokenizer's standard non-thinking chat template (empty closed think section in the input, never forced output tokens). No logit masks, candidate renormalization, constrained decoder or target labels in input. Choose the first candidate in the listed order satisfying >=6/8 correct codes and >=7/8 strictly bare outputs. Run all three for auditability. Choose only by baseline capability, format and tokenizer boundaries, never intervention success. If none passes, stop with a task-design limitation. Freeze its paired modest paraphrase before intervention development and never retune it.

Canonical continuation is the exact uppercase ISO code with no prefixed space. Audit concatenation boundaries for each formatted input; use full code tokenization, not convenient substrings. Record source and target token sequences, shared prefix and first differing token. Multi-token ISO codes often begin with country-related letters. First-token results concern a currency-code prefix, not complete currency identification; distinguish a shift before any generated answer prefix from changes after teacher forcing one. A shared code prefix makes first-token target/source log odds identically zero. Its first differing token is conditional on that prefix and cannot be treated as completely cue-free evidence.

## Development and selection

Grid is exactly 3 post-block indices {9,13,17} × positions {final prompt token, final city-span token} × per-position relative caps {0.01,0.03,0.05} = 18 configs. The source strength-one least-squares coordinate-swap delta is scaled DOWN when necessary; never amplified to fill the cap. Qwen final learned RMSNorm weights enter both J and ordinary directions. J columns: ((W[country ids] * g) @ J[layer]).T; ordinary columns: (W[country ids] * g).T. No individual column normalization; use the pseudoinverse with nonorthogonal vectors. Opposite negates the bounded direction; it is not a mathematical inverse.

Choose configurations by mean paired full-code target/source log-odds gain on the main-template development cases. Eligibility: no more than one extra malformed-format output versus baseline and no cap/matching failures. Negative effects remain eligible; no positive-effect gate. Scores within 0.10 nats of the best are effectively tied: prefer smaller cap, then earlier block, then final-city location. Save all generations and numerical results. If no configuration meets instrumentation/format eligibility, report that limitation instead of relaxing the rule.

All evaluation data, aliases, position definitions, main/paraphrase templates, selected settings, code, scoring and statistics hashes are frozen before evaluation. Evaluation is 32 cities × two templates × 10 conditions = 640 runs: baseline; instrumented no-op; J-target; opposite; unrelated table/chair; ordinary-country geometry; three random draws; old block-21/all-prompt/strength-4 bridge. Bridge is uncapped and excluded from H1 success; it uses the original intervention unchanged except prompt positions/inputs.

## Precision and matched controls

Model residuals are bf16. Calculations use float32 MPS; small pseudoinverses are float64 CPU as in the first study. Enforce actual norm AFTER rounding each changed vector to bf16, with absolute relative-cap tolerance 1e-5 (0.001 percentage points). Target fitting monotonically shrinks the proposed delta until the applied norm meets the cap. A zero/small delta is never scaled up.

Each control uses the target's actual applied norm at the identical block/position. Distinct deterministic random draws per city and repeat; reuse across main/paraphrase for paired wording comparison. Normalize a control direction and numerically adjust its scalar under bf16 rounding to match within max(2% of target norm, 1e-5 times residual norm), while respecting the same actual cap. This tolerance is not permission to exceed the cap. Search the scalar for magnitude only, never model outcomes. If bf16 granularity prevents a match, record an instrumentation failure and stop evaluation for diagnosis, not scientific interpretation. For target and controls save requested, actual, relative norms, scale, matching error and original/changed vector arrays. No-op is bit-exact.

## Scoring, outcomes and checks

Greedy generation to EOS or 64 new tokens, no custom content stop. Save complete produced text/IDs including special termination IDs and explicitly mark capped outputs as truncated. Save full unmasked first-answer logits and softmax probabilities, not just top-k. Primary full-code likelihood is SUM of all source or target code token log probabilities (no restricted softmax). Scores include no answer-specific prefix before the first token. Also calculate first-token log odds directly from the original first distribution; first-differing-token log odds if a common prefix exists; later contribution = full-code minus first-token log odds. Use explicit incremental cached teacher forcing from the same intervened prefill; no state shared between separate runs. Save per-token probabilities and boundaries.

Format rule: after removing only tokenizer special tokens, Unicode NFKC normalization and outer whitespace, the ENTIRE generated text must be exactly one valid uppercase three-letter current ISO code. Trailing explanation, punctuation, code fences or lowercase are format failures. Exact code knowledge and clean format are separate labels. Detect currency codes anywhere as whole words, preserving prefix/suffix text; source/target/other first currency-code classifications are separate from malformed/truncated flags. Country/nationality detection uses fixed labels for all 24 countries plus common variants; currency names (including Unicode đồng, złoty) are secondary markers for substantive-prefix audit, never scored as clean ISO success. Any non-whitespace content before an emitted code is a substantive prefix; country/nationality flags are a subset, with code-internal country letters reported separately. Literal country words in input are not supplied; intended city characters are located by tokenizer offset mapping.

Verify no-op logits/probabilities/tokens, city/final token positions, prompt-only hook calls, fresh caches, cached/uncached decoding on representative development cases, code boundaries and explicit joint-likelihood factorization, cap enforcement/control matching, Unicode and malformed examples. Record nonzero checks and failures. Unsupported/cap-failed runs are instrumentation findings, never null causal effects.

## Analysis, uncertainty and representative cases

All 32 cases, and within-template baseline-correct subsets, reported. Primary estimate averages the two template effects within each city; uncertainty uses 16 source-country clusters, 10,000 paired bootstrap resamples, seed 20260908. Report each template separately and paired paraphrase-minus-main effects; templates are repeats, not independent questions. Average random controls within city/template before inference. Report full likelihood changes, first-token changes, later-token changes, J-minus-ordinary/random/unrelated/opposite comparisons, clean target counts, code-with-prefix counts, malformed/truncated rates, cap distributions and all cases. Intervals do not remove dependence from reused countries, fixed cyclic targets, shared templates or feature directions. No claims about broad population rates.

Compelling example: among baseline-correct clean target successes, largest first-token log-odds gain averaged across templates; if none, largest first-token gain and explicitly label absence of a behavioral success. Counterexample: baseline-correct case with the smallest primary full-code gain, emphasizing failed predicted target if available. Retain every case in machine-readable and tabular outputs. No new explanation requests or evaluation-time retuning.
