---
title: "ESC31 Gene-tagging Programme Resilience Analysis"
format:
  html:
    toc: true
    page-layout: full
    theme: default
    lightbox: true
    mainfont: system-ui
    css: esc31.css
    favicon: favicon.ico
    include-in-header:
      text: |
        <link rel="icon" href="favicon.ico" sizes="any">
        <style>
        .mobile-scroll-table {
          max-width: 100%;
          overflow-x: auto;
          -webkit-overflow-scrolling: touch;
        }
        .esc31-code-tools {
          display: none;
        }
        </style>
    include-before-body: nav.html
    embed-resources: true
execute:
  enabled: false
bibliography: references.bib
link-citations: true
---

::: {.callout-important}
## Continue annual gene tagging wherever possible

**GT skipping is not an operational recommendation from this analysis.** The
2027--2029 proposal specifies an annual programme of approximately 5,000 age-2
releases, 13,000 age-3 harvest samples, and a target CV of 0.25
[@PreeceEveson2026GTProposal]. That annual full-effort programme is the
reference case on this page. Deliberate gaps and an unplanned failure are
counterfactual resilience tests: they show what could happen if monitoring is
reduced or disrupted, not that alternate-year GT is acceptable.
:::

## Purpose

The analysis examines the value and resilience of the annual GT programme and
the scientific and management consequences of reduced effort or missing field
years. Two consequences need to be kept separate:

1. **Stock assessment:** less information about individual recruiting cohorts
   may alter estimated recruitment, stock status, uncertainty, and the
   assessment posterior used to condition future operating models.
2. **Projection and management procedure:** missing recent recruitment
   estimates may alter the GT component of the Cape Town Procedure (CTP), TAC
   paths, and short- and long-term risk.

The objective is to compare monitoring programmes transparently while keeping
the accepted assessment and the recommended annual programme unchanged. Any
proposal to alter GT frequency would remain an exceptional circumstance and
would require the formal MSE, MP-review, and meta-rule process.

## Evidence and previous work

### Formal work already completed

The 2025 review of changes to the GT program
([official PDF](https://www.ccsbt.org/system/files/2025-08/ESC30_20_AU_Implications_of_changes_to_GT_program.pdf))
considered three distinct cost-saving mechanisms
[@PreeceEtAl2025GTChanges]:

- less frequent tagging, specifically every second year or two years out of
  three;
- smaller release and harvest sample sizes, with lower precision; and
- fewer fieldwork days, with a higher chance of failing to obtain enough
  releases.

These are not interchangeable experiments. Skipping a complete year creates a
missing abundance estimate; reducing effort usually retains an estimate but
makes it less precise and increases the risk of complete failure.

The same review explains why recruitment information matters. GT is currently
the only fishery-independent absolute abundance estimate for age-2 fish used in
the assessment and CTP. Without an informative recruitment index, a poor cohort
may not be clearly detected until it affects the spawning stock roughly
10--15 or more years later. The earlier aerial-survey work is also relevant:
MSE testing following its cancellation indicated delayed rebuilding and lower
average catch without a recruitment index [@PreeceEtAl2025GTChanges].

For the current CTP, the recent GT abundance estimates are combined using
numbers of matches as weights. A skipped year therefore removes information
from the recent window rather than being replaced with an average observation.
The 2025 review anticipated greater stochastic variation in TAC advice when
there are gaps, and noted that maintaining the same risk criteria during MP
retuning would generally reduce mean TAC.

There is also one useful observed precedent. Cancellation of tagging in 2020
left the 2018 **birth cohort**--which would have been released at age 2 in
2020--without a GT abundance estimate. The CTP was designed to
continue with an occasional missing GT value, but the 2023 assessment
discussion noted that the estimated strength of this cohort was instead
informed mainly by the high 2022 CPUE value and catches of ages 3 and 4. This
does not demonstrate a general GT-by-CPUE effect, but it explains why a
separate CPUE-q interaction is scientifically relevant.

OMMP16 subsequently placed the following item directly in the 2026 projection
workplan: allow GT years to be skipped in projections and use the projection
framework to explore short- and long-term risks
([local PDF](../OMMP16/Report_of_OMMP16.pdf), paragraph 54)
[@CCSBT2026OMMP16].

### 2026 internal context

The subsequent 2027--2029 proposal resolves the immediate programme reference:
GT is proposed annually, with about 5,000 releases and 13,000 harvest samples
per cohort and a target CV of 0.25 [@PreeceEveson2026GTProposal]. The 2025 and
2026 cohorts are already in the pipeline. The 2025 cohort has 3,680 releases
and 17,000 collected harvest samples; the 2026 cohort has 3,066 releases, with
its 13,000-sample harvest target due in 2027
[@PreeceEtAl2026GTUpdate; @PreeceEveson2026GTProposal]. These
observed and committed quantities replace the generic future sample sizes in
the long-horizon comparison below.

Email discussion in April and May 2026 did not define a quantitative terms of
reference. It identified two candidate forms of reduction--a smaller annual
exercise and a lower survey frequency--and asked how either would affect the
robustness of ESC advice. It also recommended separating this work from the
communication of the new MCMC-based risk results and using 2026, if necessary,
to make the code ready for the fuller MP review.

The 2025 paper's then-current schedule included no tagging in 2026. That
schedule should not be assumed here: the April 2026 correspondence states that
the full 2026 exercise had been reinstated. Scenario years therefore remain to
be specified from the latest agreed program rather than copied from the 2025
paper.

The July 2026 discussion with Ana adds two useful design cautions:

- exploratory results could help internal planning for next year even if they
  are not presented at ESC31; and
- past work found it difficult to show a large effect from GT gaps alone and
  obtained a stronger effect only when GT gaps were combined with failing CPUE,
  represented as a change in catchability q.

The first comparison should therefore isolate GT frequency under the standard
CPUE assumptions. Any CPUE-q failure should be a clearly labelled, secondary
joint-stress experiment, not part of the definition of the GT-only effect.

## Define the calendar before defining scenarios

Three different years can be attached to one GT abundance estimate:
@tbl-gtskip-calendar defines the terms used throughout this page.

::: {.mobile-scroll-table}
| Term | Current data field | Meaning |
|---|---|---|
| Tagging or release year | `RelYear` | Age-2 fish are released in this field year. |
| Harvest sampling or recapture year | `RecYear` | The tagged cohort is sampled, currently represented as `RelYear + 1`. |
| Data-availability year | CTP schedule | The completed estimate becomes available to a later TAC calculation. |

: Calendar terms used for gene-tagging release, recapture, and data availability. {#tbl-gtskip-calendar}
:::

Every scenario should be stored and reported in all three year conventions.
This will prevent a nominal "2028 gap" from meaning a 2028 field season in one
place and a 2028 data exchange in another.

::: {.callout-important}
## Projection calendar correction complete

The `sbt` projection interface now distinguishes `gt_skip_years`, which names
release years, from `gt_skip_recapture_years`, which independently names
harvest-sampling or recapture years. The generated template validates the
release and recapture calendars and retains the one-year relationship
`RecYear = RelYear + 1` for rows that remain. Unit tests cover release-only,
recapture-only, and combined omissions, including the fact that a literal
shutdown of odd-calendar-year fieldwork can remove every otherwise retained
even-year release under the current one-year lag.
:::

## Experiment A: consequences for the stock assessment

The first assessment experiment can use historical data omission as a
controlled sensitivity analysis:

1. retain the accepted ESC31 assessment as the baseline;
2. remove GT rows according to agreed release-year patterns;
3. refit the otherwise unchanged assessment;
4. compare the posterior displacement and loss of precision; and
5. project from each accepted posterior using the same future assumptions.

Suggested assessment summaries are:

- recruitment by cohort, especially cohorts whose GT estimate was withheld;
- spawning biomass, total reproductive output, depletion, and rebuilding
  probabilities;
- recruitment variability and other parameters that GT data may inform;
- changes in the posterior uncertainty carried into projections;
- GT likelihood and predictive diagnostics; and
- optimizer and MCMC diagnostics for every refit.

This retrospective exercise cannot estimate statistical bias because the true
historical population is unknown. It measures sensitivity and loss of
information. Bias, interval coverage, and repeated-assessment performance need
simulated operating models with known truth.

## Experiment B: consequences within projections and the CTP

For each accepted starting posterior draw, the projection experiment should:

1. simulate the same underlying population and non-GT monitoring processes;
2. apply an explicit release-year schedule to the simulated GT observations;
3. pass the resulting missing estimates into the existing CTP update schedule;
4. retain complete CTP diagnostics and TAC paths; and
5. compare scenarios using paired posterior draws and, as far as possible,
   common random numbers.

Primary summaries should include:

- the GT estimates and numbers of matches actually available at every CTP
  calculation;
- the number and age of usable GT estimates in the CTP's recent window;
- TAC level, TAC change, TAC variability, and cumulative catch;
- short- and long-term stock risks using the same definitions as the main
  projection report;
- the delay in management response to weak recruitment; and
- the frequency of CTP calculations with too little usable GT information.

The common-random-number safeguards are implemented. Projection simulations use
component-specific random streams for CPUE, GT, HSP, and POP, and GT matches
are keyed so that a retained release year receives the same draw whether or not
other release years are omitted. Unit tests confirm that removing GT rows does
not shift the HSP or POP draws. A paired comparison must still use the same
posterior-row selection, component seeds, and non-GT inputs in both arms.

## Programme-comparison scenarios

The four scenarios in @tbl-gtskip-candidate-scenarios use the same accepted
historical posterior. They differ only in future GT effort or availability.
The annual full-effort scenario is the proposed programme; the other three are
counterfactual precision and resilience tests.

::: {.mobile-scroll-table}
| ID | GT design | Purpose |
|---|---|---|
| GT0 | Annual full effort: 5,000 releases and 13,000 harvest samples from 2027 | Proposed 2027--2029 programme and reference case. |
| GT1 | Annual reduced effort: 5,000 releases and 10,000 harvest samples from 2027 | Isolate the documented lower harvest-sample design while retaining annual estimates. |
| GT2 | Retain the 2027 release, then deliberately omit alternate release cohorts from 2028 | Counterfactual test of regular gaps; not an operational recommendation. |
| GT3 | GT2 plus unplanned failure of the 2029 release cohort | Stress test three consecutive missing cohort estimates (2028--2030); not a proposed programme. |

: Gene-tagging scenarios in the common-posterior long-horizon comparison. Only GT0 describes the proposed programme. {#tbl-gtskip-candidate-scenarios}
:::

The formal proposal covers 2027--2029. For a controlled long-horizon
comparison, each scenario's 2027 design is held constant after 2029 through
the end of the simulated programme. That continuation is an experimental
assumption, not an approved fieldwork schedule.

The deliberate-skip arm omits release cohorts **2028, 2030, ..., 2054** and
their following-year recapture work. The combined stress arm makes those same
deliberate omissions and adds an unplanned failure of the **2029** release,
leaving the 2028--2030 cohort estimates consecutively unavailable.

All four use the standard CPUE process. A CPUE-q change is not mixed into this
comparison; any joint-stress test would need a separately specified start
year, size, duration, and biological rationale.

## What the current projection code can do

@tbl-gtskip-capabilities distinguishes implemented capabilities from the
remaining extensions needed for a broader program-design analysis.

::: {.mobile-scroll-table}
| Capability | Current status | Consequence for this experiment |
|---|---|---|
| Represent a missing future GT estimate in the CTP | **Implemented, tested, and exercised in both the 100-draw screen and 2,000-draw long-horizon comparison below.** Separate arguments omit release years and recapture years. A missing release-year estimate becomes `gtN = NA` and `gtR = 0`; the CTP weighted mean uses the remaining finite estimates with positive matches. | The independent production review gate passed for all four programme scenarios. |
| Apply the OMMP16-style GT availability lag | **Implemented and verified.** The default CTP schedule uses a four-year lag from TAC implementation year to the latest GT release year. | The release, recapture, data-availability, and nine CTP implementation years were checked independently for every scenario. |
| Detect a window with no usable GT estimate | **Implemented.** The CTP stops rather than silently inventing an estimate. | Long or clustered gaps need deliberate failure-handling rules for a full MSE. |
| Apply one reduced release and harvest sample size to all future GT years | **Implemented.** | Supports a constant lower-effort sensitivity. |
| Apply different effort levels in different future years | **Implemented and tested.** Release and harvest sample-size vectors are keyed by release year. | The observed 2025--2026 pipeline is shared by all four long-horizon arms before the 2027 programme scenarios diverge. |
| Simulate a future CPUE-q failure or step change | **Not exposed as a projection scenario.** The current path carries forward fitted catchability and creep. | The optional joint-stress experiment needs a defined q trajectory and code support. |
| Refit the full stock assessment at future assessment years | **Not implemented in `run_projections()`.** It projects from a fixed accepted posterior and runs the CTP's internal model, but does not update the full assessment posterior. | Experiment A must use separate refits. A fully closed-loop assessment experiment would be a larger MSE extension. |
| Keep all non-GT stochastic draws paired when GT rows are removed | **Implemented and tested.** CPUE, GT, HSP, and POP use component-specific streams, and retained GT rows use stable release-year keys. | Use the same posterior rows, component seeds, and non-GT inputs in both arms. |
| Retain weak-stock draws when nominal TAC cannot be taken | **Implemented, tested, and independently audited.** The opt-in operating-model rule scales every within-year allocation proportionally to keep raw harvest at or below 0.9, while retaining nominal advice and realized catch separately. The package default remains fail-closed. | Report the probability and magnitude of catch shortfall; do not hide infeasible advice by dropping draws, and interpret stock risk alongside the lower realized removals. |
| Project far enough to see spawning-stock consequences | **Implemented in the long-horizon comparison.** Dynamics extend through 2055 and the population state through 2056. | This spans nine CTP updates and more than the documented 10--15+ year pathway from an age-2 cohort to spawning-stock outcomes. |

: Current implementation capability for assessment and projection gene-tagging skip experiments. {#tbl-gtskip-capabilities}
:::

The code now contains the calendar, year-specific effort, common-posterior,
and random-stream machinery for the four-arm long-horizon comparison. This is
still an assessment-posterior-conditioned closed-loop CTP analysis, not a
formal revalidation or retuning of the MP across the full CCSBT operating-model
reference set.

## Proposed stages

### Stage 0: agree the question

- Decide whether the immediate question is CTP robustness, value to the stock
  assessment, a cost comparison, or preparation for a revised MP.
- Confirm the latest operational GT schedule and map release years to data
  availability.
- Agree whether results are internal screening material or formal ESC output.

### Stage 1: make the projection comparison identifiable

- [x] Correct and test release-year and recapture-year skipping.
- [x] Add a scenario table containing release, recapture, and availability years.
- [x] Isolate random streams for CPUE, GT, HSP, and POP observations.
- [x] Add explicit audit summaries of the GT values used at every CTP calculation.

### Stage 2: small diagnostic screen

- [x] Run a small paired subset of posterior draws for GT0 and the selected
  alternate-release-cohort scenario.
- [x] Check that the intended rows are absent and that all other stochastic inputs
  are paired.
- [x] Review CTP gradients, Hessians, biological states, and failure rates.
- Use this screen to decide whether to retain the scenario and refine the horizon and
  performance measures.

### Stage 3: assessment sensitivity

- [x] Refit the selected historical omission pattern.
- [x] Review MLE and posterior diagnostics before comparing stock quantities.
- [x] Carry the accepted alternative posterior into an otherwise matched
  100-draw projection screen.

### Stage 4: production analysis or full MSE

- [x] Run 2,000 common-posterior draws through 2055 and validate all four arms.
- [x] Independently audit the calendars, paired stochastic inputs, biological
  states, CTP optimisations, nominal advice, realized catch, stock status, and
  risk summaries.
- Add the CPUE-q joint stress only if it has a defensible specification.
- If the question becomes a change to the CTP or long-term program design,
  move from screening projections to a formal MSE and MP-review process.

## Questions outside this comparison

- The analysis does not attach monetary costs to the four scenarios.
- A formal MP review must decide which operating-model reference set, low-
  recruitment stresses, performance measures, and tuning criteria are needed
  for full MSE revalidation.
- A CPUE-q joint stress remains separate unless evidence defines its magnitude,
  start year, and duration.
- Any proposed programme change still requires scientific review through the
  exceptional-circumstances and meta-rule process.

## Working provenance

- OMMP16 report, projection workplan paragraph 54
  [@CCSBT2026OMMP16].
- Preece, Davies, Galeano, Hillary, and Eveson (2025), especially sections
  3--6 [@PreeceEtAl2025GTChanges].
- Internal email thread, *Exploring Gene Tagging Alternatives*, 29 April--13
  May 2026.
- Internal email thread, *2026 MP evaluation - need a breakdown to justify
  potential budget*, August 2025.
- Darcy Webber--Ana Parma discussion, 23--24 July 2026.

The correspondence entries document planning context only. Formal scientific
claims and any future terms of reference should be tied to approved CCSBT
documents.

## Earlier historical-omission comparison

The earlier run retains the accepted ESC31 base fit as the reference and changes
GT availability in both the alternative historical assessment and its future
monitoring schedule. The base MLE was not refitted. The first gate was an
alternative MLE using the same model configuration, priors, parameter map,
optimizer, and biological and numerical acceptance criteria. That MLE, its
posterior, and the paired diagnostic projection are now complete.

The posterior review below uses all 3,000 retained draws from each fit. The
initial projection comparison uses 100 chain-balanced posterior draws from
each fit. It is retained as a diagnostic of historical information loss, not
as the programme recommendation or the main future-monitoring comparison. Its
different historical posteriors are specifically avoided in the common-
posterior long-horizon analysis below.

### Selected alternate-cohort design

The historical-omission experiment has one implementation:
`odd_release_cohorts`. It drops
the complete GT estimate for releases in 2017, 2019, 2021, and 2023,
including the associated recapture work one year later. It retains the four
GT estimates from releases in 2016, 2018, 2022, and 2024.

This is the requested every-second-estimate comparison. Its operational
calendar is important: retained even-year releases are recaptured in the
following odd calendar year. It therefore does **not** represent a complete
shutdown of all fieldwork in odd calendar years. A literal odd-calendar-year
shutdown would remove the recaptures for every retained even-year release and,
with the current one-year release--recapture lag, leave no usable GT
estimates.

### Reproducible run contract

- Fit artifacts are written under `ESC31/runs/gtskip/`; failed MCMC attempts
  and their diagnostics are retained rather than overwritten invisibly.
- The first review gate stops after the alternative MLE; no MCMC or projection
  is launched until that fit is accepted for further work.
- The alternative MLE must have convergence code zero, maximum gradient no
  greater than 0.01, estimability, valid biological states, exact catch
  accounting, and the accepted LL4 configuration.
- The alternative MCMC uses four chains, 150 warmup and 750 retained
  iterations per chain, dense metric, the MLE mode as the chain start,
  `adapt_delta = 0.999`, maximum treedepth 13, and seed 73015.
- Posterior acceptance requires maximum rank-normalized
  R-hat below 1.01, bulk and tail ESS at least 400, no divergences, no
  maximum-treedepth hits, and valid biological states for every retained draw.
- The base and alternative 100-draw projections use the same retained
  iteration numbers in each chain and common component-specific random
  streams. CPUE, GT, HSP, and POP streams are isolated so removing GT years
  cannot shift the HSP or POP random draws.
- Projected skip controls distinguish release years from harvest-sampling or
  recapture years. The generated calendar and the GT rows reaching each CTP
  update will be saved and reported.
- The fit, MCMC, and paired projection artifacts are retained separately from
  the accepted production assessment and projection caches.

Only files carrying the `odd_release_cohorts` experiment name are inputs or
outputs of the historical-omission analysis documented here. The older
unprefixed `gtskip_mle_*` files and
`esc31_gtskip_odd_calendar_fieldwork.sbt.rds` are preserved exploratory
artifacts for different designs; they are not accepted results and must not be
substituted for the named files. The local artifact registry in
`runs/gtskip/README.md` records that distinction.

## MLE review

::: {.callout-important}
## First MLE gate complete

The alternate-release-cohort alternative has been fitted to maximum likelihood
and passes every numerical, estimability, biological-state, catch-accounting,
harvest-wall, and LL4 gate. Its posterior and paired diagnostic projection are
reviewed separately below.
:::

The executable comparison retains four of the eight historical GT
observations. An independent input audit confirms that the numerical contents
of all other model data, the parameter map, priors, bounds, and optimizer
controls match the accepted base. @tbl-gtskip-mle-diagnostics reports the
acceptance checks. The alternative was initialized from the base MLE, but every
active parameter was then re-estimated.

::: {.mobile-scroll-table}
| Check | Requirement | Base | Alternate-cohort GT MLE |
|---|---:|---:|---:|
| Historical GT observations | 8 versus 4 | 8 | 4 |
| Final negative log-likelihood | Finite | 7616.637822 | 7600.761697 |
| Optimizer convergence code | 0 | 0 | 0 |
| Maximum absolute gradient | ≤ 0.01 | 2.60 × 10⁻¹⁰ | 5.72 × 10⁻⁵ |
| Estimability | Pass | Pass | Pass |
| Biological-state gate | Pass | Pass | Pass |
| Maximum raw seasonal harvest at age | ≤ 0.9 | 0.617533 | 0.609524 |
| Minimum populated numbers-at-age | > 0 | 1696.516 | 1723.101 |
| Maximum scaled catch error | ≤ 10⁻⁸ | 2.53 × 10⁻¹⁶ | 3.65 × 10⁻¹⁶ |
| Harvest-wall objective contribution | Report (not a gate) | 2.57 × 10⁻²¹ | 5.17 × 10⁻²² |
| Continuation contribution | ≤ 10⁻⁸ | 0 | 0 |
| LL4 standard-fleet contract | Pass | Pass | Pass |

: Numerical and biological acceptance checks for the base and alternate-cohort maximum-likelihood fits. {#tbl-gtskip-mle-diagnostics}
:::

The alternative objective is 15.876 units lower, but this is **not** evidence
of a better-fitting model: its objective omits the four GT likelihood
contributions from the omitted release cohorts. The useful comparison is
the change in fitted stock quantities in @tbl-gtskip-mle-quantities, not the
raw objective difference.

::: {.mobile-scroll-table}
| MLE quantity | Base | Alternate-cohort GT | Difference | Relative difference |
|---|---:|---:|---:|---:|
| B₀ | 6,422,189 | 6,426,113 | +3,924 | +0.06% |
| 2025 TRO | 1,653,497 | 1,777,381 | +123,884 | +7.49% |
| 2025 relative TRO | 0.2575 | 0.2766 | +0.0191 | +7.43% |
| 2025 recruitment | 3,216,124 | 3,405,529 | +189,405 | +5.89% |
| 2026 TRO | 1,717,909 | 1,860,853 | +142,944 | +8.32% |
| 2026 relative TRO | 0.2675 | 0.2896 | +0.0221 | +8.25% |
| 2026 recruitment | 3,645,944 | 3,878,668 | +232,723 | +6.38% |
| M₀ | 0.36574 | 0.36969 | +0.00395 | +1.08% |
| M₄ | 0.16167 | 0.16342 | +0.00175 | +1.08% |
| M₁₀ | 0.10736 | 0.10852 | +0.00116 | +1.08% |
| M₃₀ | 0.45771 | 0.45994 | +0.00223 | +0.49% |
| CPUE q | 0.95749 | 0.96826 | +0.01078 | +1.13% |

: Maximum-likelihood stock and parameter quantities for the base and alternate-cohort gene-tagging fits. {#tbl-gtskip-mle-quantities}
:::

![Base and alternate-cohort maximum-likelihood trajectories from 2000 through the
start-of-2026 model state. Points show every annual estimate. These are fitted
MLE trajectories and contain no posterior uncertainty.](runs/gtskip/gtskip_odd_release_cohorts_mle_trajectories.png){#fig-gtskip-mle-trajectories}

Retaining every second GT estimate still raises the recent fitted recruitment
peaks and the terminal stock trajectory in this MLE
(@fig-gtskip-mle-trajectories). The largest proportional
recruitment change over the fitted trajectory is about 38.9% in 2015; by 2026,
recruitment is 6.4% higher and relative TRO is 0.0221 higher. The posterior
review below evaluates uncertainty around the assessment trajectories, and the
100-draw screen then evaluates the first CTP response. The MLE result alone is
not a management conclusion. The accepted ESC31 base, grid, and production
projections remain closed while this separate robustness analysis continues.

The accepted fit is
`runs/gtskip/esc31_gtskip_odd_release_cohorts.sbt.rds`. Reproducible review
tables and the plotted trajectories are stored beside it as
`gtskip_odd_release_cohorts_mle_diagnostics.csv`,
`gtskip_odd_release_cohorts_mle_management.csv`, and
`gtskip_odd_release_cohorts_mle_review.rds`.

## Alternate-cohort posterior review

::: {.callout-important}
## Posterior gate complete

The local four-chain alternate-cohort MCMC passed every numerical and
biological-state acceptance gate. It is an accepted assessment sensitivity
posterior for this screen, not a GT-skipping projection or management
procedure result.
:::

The sampler used the pre-specified dense-metric contract: four chains, 150
warmup iterations and 750 retained iterations per chain, exact MLE-mode starts,
`adapt_delta = 0.999`, maximum treedepth 13, and seed 73015.
@tbl-gtskip-mcmc-diagnostics records the resulting diagnostics. The initial
automatic step-size search had a few warmup-only divergences; the 3,000
retained draws used for inference had none.

::: {.mobile-scroll-table}
| Check | Requirement | Alternate-cohort posterior |
|---|---:|---:|
| Chains | 4 | 4 |
| Warmup per chain | 150 | 150 |
| Retained per chain | 750 | 750 |
| Total retained draws | 3,000 | 3,000 |
| Maximum rank-normalized R-hat | < 1.01 | 1.0076 |
| Minimum bulk ESS | ≥ 400 | 1,632.1 |
| Minimum tail ESS | ≥ 400 | 1,153.3 |
| Divergences after warmup | 0 | 0 |
| Maximum-treedepth hits | 0 | 0 |
| State draws expected / checked | 3,000 / 3,000 | 3,000 / 3,000 |
| Invalid or non-finite state draws | 0 | 0 |
| Maximum raw seasonal harvest at age | ≤ 0.9 | 0.8731 |
| Minimum populated numbers-at-age | > 0 | 603.348 |
| Maximum continuation penalty | ≤ 10⁻⁸ | 0 |
| Overall posterior gate | Pass | Pass |

: Sampling and biological-state diagnostics for the accepted alternate-cohort gene-tagging posterior. {#tbl-gtskip-mcmc-diagnostics}
:::

### Historical GT rows

The alternate fit skipped releases in **2017, 2019, 2021, and 2023** and their
associated recapture years **2018, 2020, 2022, and 2024**. It retained releases
in 2016, 2018, 2022, and 2024, recaptured respectively in 2017, 2019, 2023,
and 2025. This is the linked cohort-level schedule in
@tbl-gtskip-historical-data; it is not a blanket shutdown of all work in odd
calendar years.

There is no 2020 release row in @tbl-gtskip-historical-data because tagging
was cancelled in that field year. That source-data gap is separate from the
retained/omitted pattern imposed by this sensitivity.

::: {.mobile-scroll-table}
| Release year | Release age | Recapture year | Releases | Scanned samples | Matches | Alternate-fit status |
|---:|---:|---:|---:|---:|---:|---|
| 2016 | 2 | 2017 | 2,952 | 15,389 | 20 | Retained |
| 2017 | 2 | 2018 | 6,480 | 11,932 | 67 | Omitted |
| 2018 | 2 | 2019 | 6,295 | 11,980 | 66 | Retained |
| 2019 | 2 | 2020 | 4,242 | 11,109 | 31 | Omitted |
| 2021 | 2 | 2022 | 6,401 | 10,742 | 41 | Omitted |
| 2022 | 2 | 2023 | 5,084 | 14,714 | 38 | Retained |
| 2023 | 2 | 2024 | 2,759 | 13,297 | 14 | Omitted |
| 2024 | 2 | 2025 | 3,522 | 11,011 | 11 | Retained |

: Historical gene-tagging data and the rows retained or omitted in the alternate-cohort assessment fit. {#tbl-gtskip-historical-data}
:::

### Posterior trajectories and GT fits

@fig-gtskip-posterior-trajectories compares the assessment trajectories using
all 3,000 retained draws from each posterior. These are separately sampled
base and alternate assessment distributions; they are not paired projection
draws.

![Base and alternate-cohort posterior trajectories from 2000 through the
start-of-2026 model state. Lines are posterior medians and ribbons are
equal-tailed 95% credible intervals from all 3,000 retained draws in each
fit.](runs/gtskip/gtskip_odd_release_cohorts_mcmc_trajectories.png){#fig-gtskip-posterior-trajectories}

At 2026, median relative TRO is 0.3044 in the alternate fit versus 0.2817 in
the base, an absolute difference of 0.0227 (8.0% relative to the base median).
The corresponding 95% intervals are 0.2447--0.3829 and 0.2216--0.3526.
Median recruitment is 3.819 million versus 3.680 million, a difference of
0.139 million (3.8%); its two wide 95% intervals also overlap substantially.
The upward terminal-TRO direction seen at the MLE therefore remains in the
posterior medians, but the assessment uncertainty is material and this is not
a paired test of projected outcomes.

The fitted GT expectations in @fig-gtskip-posterior-gt-fit use the sample sizes
in @tbl-gtskip-historical-data. Orange points are observations used in the
corresponding likelihood. Grey crosses locate the four observations omitted
from the alternate fit; they are shown only for context and did not contribute
to its likelihood.

![Observed and posterior fitted historical GT match counts for the base and
alternate-cohort assessments. Blue open points are posterior medians and
vertical blue intervals are equal-tailed 95% credible intervals for expected
matches. Grey crosses in the alternate panel are omitted observations, not
fitted data.](runs/gtskip/gtskip_odd_release_cohorts_mcmc_gt_fit.png){#fig-gtskip-posterior-gt-fit}

## Paired 100-draw CTP projection screen

::: {.callout-important}
## Diagnostic projection gate passed

The annual-GT reference and alternate-release-cohort arms each completed 100
chain-balanced draws. All 200 historical states and all 400 CTP update
optimisations across the two arms passed their configured gates. This accepts
the run as a diagnostic screen; it does not promote it to the 2,000-draw
production assessment or make it ESC advice.
:::

The comparison changes two linked parts of the workflow: the alternate arm
starts from the accepted historical-omission posterior reported above and also
omits the corresponding future GT release cohorts. It therefore represents a
continued alternate-cohort program, not a projection-only test in which both
arms start from the same posterior.

The 100 rows comprise 25 retained iterations from each of four chains. Both
arms use draw seed 44, projection seed 102, the same retained iteration
numbers, the same recruitment and selectivity rules, the same fixed TAC and
fleet allocation, and component-specific common random streams. The projection
covers model and monitoring years 2022--2035 and the start-of-2036 population
state.

### Future GT calendar

The future schedule in @tbl-gtskip-future-calendar skips releases in **2025,
2027, 2029, 2031, and 2033** and their linked recaptures in **2026, 2028, 2030,
2032, and 2034**. Retained releases still require recapture work in the
following odd calendar year.

This retained 100-draw screen is deliberately counterfactual from release
year 2025: it skips the 2025 release even though that field programme has
already occurred. It is therefore not the current operational pipeline. It
also omits **odd** release years, whereas the later common-posterior GT2 and GT3
arms retain 2027 and omit **even** release years from 2028. “Alternate cohort”
describes the frequency in both experiments, not a shared parity or calendar.

::: {.mobile-scroll-table}
| Release year | Recapture year | Alternate-arm GT estimate |
|---:|---:|---|
| 2025 | 2026 | Skipped |
| 2026 | 2027 | Retained |
| 2027 | 2028 | Skipped |
| 2028 | 2029 | Retained |
| 2029 | 2030 | Skipped |
| 2030 | 2031 | Retained |
| 2031 | 2032 | Skipped |
| 2032 | 2033 | Retained |
| 2033 | 2034 | Skipped |
| 2034 | 2035 | Retained |

: Future release and recapture calendar used in the alternate-release-cohort projection. {#tbl-gtskip-future-calendar}
:::

### Projection diagnostics

@tbl-gtskip-projection-diagnostics records the complete numerical gate. The
alternate arm's maximum CTP gradient, \(4.9811\times10^{-4}\), passes the
\(5\times10^{-4}\) tolerance but is close enough that it should be watched in
a larger run. Every final CTP Hessian was positive definite and every
historical state passed.

::: {.mobile-scroll-table}
| Check | Requirement | Annual GT | Alternate release cohorts |
|---|---:|---:|---:|
| Projection draws expected | 100 | 100 | 100 |
| Historical states checked | 100 | 100 | 100 |
| Historical states passed | 100 | 100 | 100 |
| CTP updates expected | 200 | 200 | 200 |
| CTP updates passed | 200 | 200 | 200 |
| Positive-definite CTP Hessians | 200 | 200 | 200 |
| Maximum absolute CTP gradient | ≤ 0.0005 | 0.0004668 | 0.0004981 |
| Elapsed time | Report only | 144.8 s | 144.9 s |

: Numerical and biological-state diagnostics for the paired 100-draw CTP projection screen. {#tbl-gtskip-projection-diagnostics}
:::

### CTP inputs and TAC response

The first CTP update in 2030 can use GT information through release year 2026;
the second, in 2033, can use information through 2029. @tbl-gtskip-ctp-results
shows that skipping estimates did not simply lower the weighted GT abundance.
The remaining match-weighted estimates produced a slightly higher median at
both updates, and the median GT multiplier stayed at its neutral value of one
in both arms.

::: {.mobile-scroll-table}
| CTP implementation year | Latest GT release year | Median GT abundance: annual | Median GT abundance: alternate | Median GT multiplier: annual / alternate | Median TAC: annual | Median TAC: alternate | Decrease / unchanged / increase: annual | Decrease / unchanged / increase: alternate |
|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| 2030 | 2026 | 2,244,647 | 2,288,038 | 1.00 / 1.00 | 26,647.0 | 26,647.0 | 0 / 0 / 100 | 0 / 0 / 100 |
| 2033 | 2029 | 1,953,125 | 2,222,222 | 1.00 / 1.00 | 26,802.3 | 26,818.8 | 1 / 37 / 62 | 2 / 30 / 68 |

: GT inputs and TAC outcomes at each CTP update in the 100-draw screen. Counts in the final columns are numbers of projection draws. {#tbl-gtskip-ctp-results}
:::

Figure @fig-gtskip-projection-tac shows the resulting marginal TAC paths.

![Nominal total TAC for the annual-GT and alternate-release-cohort arms. Lines
show medians and ribbons show empirical 95% intervals across 100 draws. Dashed
vertical lines mark CTP update years; TAC is fixed before
2030.](runs/gtskip/gtskip_odd_release_cohorts_projection_100_tac.png){#fig-gtskip-projection-tac}

The central TAC paths are very similar, but the result is more nuanced than
“no effect”:

- The 2026--2029 TACs are prescribed and therefore identical.
- In 2030, the median unconstrained TAC is about 35,542 t in the annual arm and
  35,502 t in the alternate arm. Every draw reaches the 3,000 t increase
  limit, so both implemented medians are exactly 26,647 t through 2032.
- At the 2033 update, the marginal medians differ by only 16.5 t
  (26,802.3 versus 26,818.8 t). Their empirical 95% intervals are
  26,647--29,647 t and 26,393--29,647 t, respectively.
- The draw-level responses are not identical: 74 of 100 paired computational
  rows differ at the 2033 update, and the empirical paired difference interval
  is -2,852 to +2,540 t. This artificial common-random-number pairing is useful
  for a screen, but it is not a credible interval from a formally joint
  posterior.
- Median cumulative nominal TAC over 2026--2035 is 253,960 t for annual GT and
  253,998 t for the alternate arm. The paired median difference is zero, while
  its empirical 95% interval is -8,557 to +7,619 t.

The first update is therefore uninformative about the operational effect of GT
gaps because the TAC-change cap binds. The second update suggests very similar
central TAC advice but enough draw-level variation that equivalence should not
be claimed from 100 draws.

### Projected stock status and risk

Figure @fig-gtskip-projection-tro and Table @tbl-gtskip-projection-tro report
the corresponding stock-status screen.

![Projected relative total reproductive output for the annual-GT and
alternate-release-cohort arms. Lines show medians and ribbons show empirical
95% intervals across 100 draws. The dashed horizontal line is the relative-TRO
risk threshold of 0.2.](runs/gtskip/gtskip_odd_release_cohorts_projection_100_relative_tro.png){#fig-gtskip-projection-tro}

::: {.mobile-scroll-table}
| State year | Annual GT median (95% interval) | Alternate median (95% interval) | Difference in marginal medians |
|---:|---:|---:|---:|
| 2030 | 0.330 (0.265--0.421) | 0.363 (0.280--0.461) | +0.034 |
| 2033 | 0.345 (0.262--0.498) | 0.387 (0.289--0.516) | +0.042 |
| 2036 | 0.344 (0.241--0.571) | 0.384 (0.270--0.546) | +0.040 |

: Selected relative-TRO summaries from the paired 100-draw projection screen. {#tbl-gtskip-projection-tro}
:::

The alternate arm retains the higher terminal stock-status direction already
seen in its historical assessment posterior. No draw in either arm falls below
relative TRO 0.2 during the configured short-term window (2030--2032) or
longer window (2033--2036). With 100 draws this means only that the observed
screening frequency was zero; it does not establish a zero risk probability.

The higher alternate trajectory must not be attributed solely to future GT
skipping. It also reflects conditioning on the different accepted historical
posterior, and the intervals overlap substantially. Moreover, the current
horizon is too short to capture the full documented 10--15+ year path from a
future age-2 cohort to spawning-stock outcomes.

### Interpretation and next gate

The screen supports three provisional comments:

1. Removing alternate GT estimates did not automatically depress the CTP's
   weighted GT input; which cohorts remain and their match weights matter.
2. The median TAC path was effectively unchanged through the first update and
   differed by only 16.5 t at the second, largely because the first update hit
   the TAC-change cap and the median GT multiplier remained neutral.
3. Individual draw responses and cumulative differences were much wider than
   the marginal medians suggest, while the stock-status comparison also
   includes historical-posterior displacement.

Those limitations motivate the common-posterior long-horizon comparison below.
The earlier result remains useful for showing that historical GT omission can
move the assessment posterior, but its 16.5 t median TAC difference is not used
to judge whether a future programme is acceptable.

## Common-posterior long-horizon closed-loop comparison

This comparison isolates future monitoring. Every arm uses the same 2,000
chain-balanced draws from the accepted base posterior, the same future
recruitment and selectivity paths, the same catch allocation and removal
assumptions, and component-specific common random numbers. Only the future GT
effort and availability schedule changes.

The projection covers dynamics from 2022 through 2055 and the resulting state
through 2056. Fixed TACs apply from 2026 through 2029, after which the CTP is
run at nine implementation years: 2030, 2033, 2036, 2039, 2042, 2045, 2048,
2051, and 2054. This removes the two-update horizon limitation of the earlier
screen and spans the 10--15+ year pathway from juvenile recruitment to the
spawning stock.

Nominal CTP TAC advice and realized catch are retained separately. No posterior
draw is discarded or replaced if its stock becomes too small to take the
nominal TAC. Instead, an explicit operating-model implementation rule reduces
all seasonal and fishery allocations in that draw-year by one common factor
until the raw combined seasonal harvest rate is no greater than 0.9. The report
therefore audits catch shortfall as an outcome. Because this rule can reduce
late-horizon removals in weak-stock draws, risk results must be interpreted
together with both nominal TAC and realized catch rather than as an
unqualified consequence of the monitoring design.

The programme's target CV of 0.25 is retained as an operational design target,
not imposed as a model acceptance gate. The results report simulated match
counts and use 16 matches (the crude Poisson quantity \(1/0.25^2\)) only as a
descriptive precision proxy. That proxy is not a replacement for a formal
design-based CV calculation.

::: {.callout-note}
## Scope of this closed-loop result

This is a closed-loop, posterior-conditioned CTP simulation: monitoring data
are generated from each operating-model path, passed to the CTP, and the
resulting TAC changes feed back into subsequent population and observation
paths. It is substantially more informative than the 100-draw screen. It is
not, however, formal revalidation or retuning of the MP across the complete
CCSBT operating-model reference set, and it cannot authorize a change from the
annual programme.
:::

::: {.callout-important}
## Production comparison complete; retain annual gene tagging

All four 2,000-draw arms passed the independent numerical and biological
review. Reducing the annual harvest sample lowered simulated match precision
but retained five recent estimates in every CTP window and produced results
very close to the full-effort reference in this experiment. Deliberate gaps
reduced the recent-information window to two or three estimates, and the added
failure reduced it to one estimate at the 2036 CTP calculation. Those gaps
increased TAC variability and gave slightly poorer long-horizon stock-risk
indicators, although the paired simulation intervals were broad and included
zero. These counterfactuals do not provide a basis for reducing the proposed
annual programme.
:::

### Independent validation

The production run reused the same 2,000 chain-balanced posterior rows and
paired non-GT stochastic inputs in all arms. Each arm passed all 2,000
historical-state checks and all 18,000 CTP optimisations. Every final CTP
Hessian was positive definite, the maximum absolute gradient met the
pre-specified $5 \times 10^{-4}$ tolerance, populated abundance remained
positive, raw harvest reached but did not exceed the configured 0.9 ceiling
beyond its $10^{-10}$ numerical tolerance, and the feasibility penalty was
effectively zero. Table @tbl-gtprogramme-diagnostics records these gates.

::: {.mobile-scroll-table}
| Programme scenario | Historical states passed | CTP updates passed | Maximum CTP gradient | Minimum Hessian eigenvalue | Overall gate |
|---|---:|---:|---:|---:|---:|
| Annual full effort | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499976 | 12.025 | Pass |
| Annual reduced effort | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499976 | 11.897 | Pass |
| Deliberate alternate-cohort skips | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499930 | 12.051 | Pass |
| Deliberate skips plus unplanned failure | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499953 | 12.034 | Pass |

: Independent numerical and biological review of the four long-horizon programme arms. {#tbl-gtprogramme-diagnostics}
:::

### Monitoring coverage and simulated precision

Figure @fig-gtprogramme-coverage and Table @tbl-gtprogramme-monitoring separate
the number of usable cohort estimates from the precision of those retained.

![Median number of usable GT estimates in the CTP's five-cohort window. Annual
full and annual reduced effort retain all five estimates. Deliberate skips
alternate between two and three; adding the unplanned 2029 failure leaves only
one usable estimate at the 2036 calculation.](runs/gtskip/gtprogramme_2000_through_2055_ctp_coverage.png){#fig-gtprogramme-coverage}

Reducing annual harvest sampling from 13,000 to 10,000 lowered the median
simulated matches per active cohort from 29 to 22 and more than doubled the
mean frequency below the descriptive 16-match precision proxy. It did not
remove an estimate from the CTP window. The skip scenarios retained the
full-effort precision for cohorts that were actually sampled, but had far
fewer cohort estimates available. Match precision for an observed cohort and
coverage of recent cohorts are therefore distinct consequences.

::: {.mobile-scroll-table}
| Programme scenario | Median matches per active cohort | Probability below 16-match proxy | Minimum median usable estimates per CTP window |
|---|---:|---:|---:|
| Annual full effort | 29 | 10.9% | 5 |
| Annual reduced effort | 22 | 23.2% | 5 |
| Deliberate alternate-cohort skips | 29 | 10.9% | 2 |
| Deliberate skips plus unplanned failure | 29 | 10.9% | 1 |

: Simulated precision of available future GT estimates and coverage of the CTP's recent five-cohort window. The 16-match quantity is a descriptive Poisson proxy for CV 0.25, not a formal design-based precision calculation. {#tbl-gtprogramme-monitoring}
:::

### TAC response

Figures @fig-gtprogramme-tac and @fig-gtprogramme-paired-tac show marginal and
paired TAC responses; Table @tbl-gtprogramme-tac-differences summarizes the
paired cumulative differences.

![Nominal total TAC under the four programme scenarios. Lines are marginal
medians and ribbons are empirical 95% simulation intervals across 2,000
draws. Dotted vertical lines mark the nine CTP implementation
years.](runs/gtskip/gtprogramme_2000_through_2055_tac.png){#fig-gtprogramme-tac}

All arms are identical through the capped 2030 response. Later marginal median
TACs are higher under deliberate gaps, especially after the added failure,
but this is not evidence that missing monitoring improves performance. Missing
cohorts change the recent match-weighted GT input in either direction rather
than imposing a downward adjustment. The median paired cumulative increase is
1,543 t with deliberate skips and 3,332 t after the added failure, only about
0.2% and 0.4%, respectively, of the annual-full median cumulative TAC of
819,398 t. Their paired 95% intervals are much wider and span zero. The higher
central nominal advice is also accompanied by slightly lower central stock
status and higher long-term risk.

![Paired nominal-TAC differences from annual full effort. Lines are mean
paired differences and ribbons are empirical paired 95% simulation intervals.
The paired median is zero at many update years: initially most draws share the
same capped response, while later non-zero differences increasingly straddle
zero even though most paired draws differ.](runs/gtskip/gtprogramme_2000_through_2055_paired_tac.png){#fig-gtprogramme-paired-tac}

::: {.mobile-scroll-table}
| Programme scenario | Median paired cumulative TAC difference (t) | Empirical paired 95% interval (t) | Mean absolute paired difference (t) | Draws with a different cumulative TAC |
|---|---:|---:|---:|---:|
| Annual reduced effort | 0 | -16,116 to 16,911 | 4,059 | 73.8% |
| Deliberate alternate-cohort skips | +1,543 | -43,030 to 52,483 | 14,348 | 86.0% |
| Deliberate skips plus unplanned failure | +3,332 | -58,170 to 72,301 | 20,413 | 88.2% |

: Paired cumulative nominal-TAC differences from the annual full-effort reference over 2026--2055. Intervals describe the simulated paired outcomes and are not posterior credible intervals for a formally revalidated MP. {#tbl-gtprogramme-tac-differences}
:::

### Stock status and risk

Figure @fig-gtprogramme-relative-tro and Table @tbl-gtprogramme-stock-risk show
the long-horizon stock distributions and the associated risk summaries.

![Projected relative total reproductive output under the four programme
scenarios. Lines are marginal medians and ribbons are empirical 95% simulation
intervals. The dashed horizontal line is the relative-TRO risk threshold of
0.2.](runs/gtskip/gtprogramme_2000_through_2055_relative_tro.png){#fig-gtprogramme-relative-tro}

The annual reduced-effort arm remains very close to full annual effort. The two
gap scenarios have progressively lower 2056 median relative TRO and higher
full-horizon probabilities of ever falling below 0.2. The central changes are
small relative to the simulation uncertainty: even the combined skip-plus-
failure arm has a paired 2056 difference interval of -0.0672 to +0.0588.
Separation appears mainly in the final years, consistent with the delay from
age-2 recruitment information to spawning-stock consequences.

::: {.mobile-scroll-table}
| Programme scenario | 2056 relative TRO, median (95% interval) | Paired median difference from annual full (95% interval) | P(2056 relative TRO < 0.2) | P(ever < 0.2), 2030--2056 |
|---|---:|---:|---:|---:|
| Annual full effort | 0.291 (0.046--0.637) | Reference | 23.9% | 26.6% |
| Annual reduced effort | 0.290 (0.047--0.640) | 0.0000 (-0.0154 to +0.0146) | 23.8% | 26.7% |
| Deliberate alternate-cohort skips | 0.287 (0.047--0.636) | -0.0011 (-0.0478 to +0.0400) | 25.3% | 28.4% |
| Deliberate skips plus unplanned failure | 0.283 (0.043--0.638) | -0.0026 (-0.0672 to +0.0588) | 26.3% | 29.4% |

: Long-horizon stock status and risk under the four programme scenarios. The paired intervals reflect common-posterior simulation differences; they are not evidence that the scenarios are equivalent. {#tbl-gtprogramme-stock-risk}
:::

In the first decade (2030--2039), the probability of ever falling below 0.2
is about 2.8% in every arm. In the final 2050--2056 window it rises from 25.6%
under annual full effort to 27.4% with deliberate skips and 28.6% with the
additional failure. Across the complete 2030--2056 horizon, the corresponding
ever-below risk rises from 26.6% to 28.4% (+1.9 percentage points) and 29.4%
(+2.9 percentage points), using the underlying unrounded probabilities (1.85
and 2.85 percentage points before rounding). The probability of being below
0.2 specifically in 2056 rises by 1.4 and 2.4 percentage points. This analysis therefore detects a
modest adverse long-term direction from missing monitoring, not a sharply
separated short-term effect; the paired final-stock intervals still span zero.

### Nominal advice and realized catch

Figure @fig-gtprogramme-catch-shortfall shows where nominal TAC advice cannot
be fully realized under the operating-model harvest ceiling.

![Probability that nominal TAC cannot be fully realized under the operating-
model harvest ceiling. The event remains rare but becomes more common late in
the projection as uncertainty in weak-stock paths expands.](runs/gtskip/gtprogramme_2000_through_2055_catch_shortfall.png){#fig-gtprogramme-catch-shortfall}

Median cumulative catch shortfall is zero in every arm, and only 2.5--2.6% of
draws experience any shortfall. The 97.5th percentile of cumulative shortfall
is about 7 t under either annual scenario, 389 t under deliberate skips, and
1,215 t after the added failure. A few weak-stock paths have much larger
annual shortfalls. Stock-risk results must therefore be interpreted alongside
realized catch: the retained weak-stock draws are not allowed to remove an
infeasible nominal TAC.

### Scientific interpretation

1. **Lower annual sample size primarily affects precision.** It retains five
   recent GT estimates in each CTP window but roughly doubles the frequency of
   active cohorts falling below the descriptive match-count proxy. In this
   isolated GT-only comparison, its TAC and stock-risk results remain close to
   annual full effort.
2. **Skipping cohorts primarily affects coverage.** Regular gaps reduce the
   five-cohort window to two or three estimates. An adjacent unplanned failure
   briefly leaves only one, increasing TAC dispersion and producing the
   largest adverse central stock-risk changes.
3. **The direction is more informative than the apparent separation.** The
   gap scenarios have higher central nominal TACs and slightly lower relative
   TRO, but paired intervals span zero and the long-horizon stock distributions
   overlap strongly. Neither equivalence nor a precise causal effect should be
   claimed from this one posterior-conditioned experiment.
4. **Annual GT remains the appropriate reference.** The comparison excludes a
   CPUE-q failure, future full-assessment refits, alternative operating-model
   reference sets, and MP retuning. Any proposed reduction still requires the
   formal MSE, MP-review, exceptional-circumstances, and meta-rule process.

The independent summaries and figures were generated with
`scripts/summarise-esc31-gtprogramme-mse.R` from the atomic four-arm comparison
and its four scenario caches. The review binds the exact input artifacts by
MD5 and stops if calendars, paired non-GT inputs, biological states, CTP
optimisations, or catch accounting fail their configured gates.
