ESC31 Gene-tagging Programme Resilience Analysis
GT skipping is not an operational recommendation from this analysis. The 2027–2029 proposal specifies an annual programme of approximately 5,000 age-2 releases, 13,000 age-3 harvest samples, and a target CV of 0.25 (Preece and Eveson 2026). That annual full-effort programme is the reference case on this page. Deliberate gaps and an unplanned failure are counterfactual resilience tests: they show what could happen if monitoring is reduced or disrupted, not that alternate-year GT is acceptable.
Purpose
The analysis examines the value and resilience of the annual GT programme and the scientific and management consequences of reduced effort or missing field years. Two consequences need to be kept separate:
- Stock assessment: less information about individual recruiting cohorts may alter estimated recruitment, stock status, uncertainty, and the assessment posterior used to condition future operating models.
- Projection and management procedure: missing recent recruitment estimates may alter the GT component of the Cape Town Procedure (CTP), TAC paths, and short- and long-term risk.
The objective is to compare monitoring programmes transparently while keeping the accepted assessment and the recommended annual programme unchanged. Any proposal to alter GT frequency would remain an exceptional circumstance and would require the formal MSE, MP-review, and meta-rule process.
Evidence and previous work
Formal work already completed
The 2025 review of changes to the GT program (official PDF) considered three distinct cost-saving mechanisms (Preece et al. 2025):
- less frequent tagging, specifically every second year or two years out of three;
- smaller release and harvest sample sizes, with lower precision; and
- fewer fieldwork days, with a higher chance of failing to obtain enough releases.
These are not interchangeable experiments. Skipping a complete year creates a missing abundance estimate; reducing effort usually retains an estimate but makes it less precise and increases the risk of complete failure.
The same review explains why recruitment information matters. GT is currently the only fishery-independent absolute abundance estimate for age-2 fish used in the assessment and CTP. Without an informative recruitment index, a poor cohort may not be clearly detected until it affects the spawning stock roughly 10–15 or more years later. The earlier aerial-survey work is also relevant: MSE testing following its cancellation indicated delayed rebuilding and lower average catch without a recruitment index (Preece et al. 2025).
For the current CTP, the recent GT abundance estimates are combined using numbers of matches as weights. A skipped year therefore removes information from the recent window rather than being replaced with an average observation. The 2025 review anticipated greater stochastic variation in TAC advice when there are gaps, and noted that maintaining the same risk criteria during MP retuning would generally reduce mean TAC.
There is also one useful observed precedent. Cancellation of tagging in 2020 left the 2018 birth cohort–which would have been released at age 2 in 2020–without a GT abundance estimate. The CTP was designed to continue with an occasional missing GT value, but the 2023 assessment discussion noted that the estimated strength of this cohort was instead informed mainly by the high 2022 CPUE value and catches of ages 3 and 4. This does not demonstrate a general GT-by-CPUE effect, but it explains why a separate CPUE-q interaction is scientifically relevant.
OMMP16 subsequently placed the following item directly in the 2026 projection workplan: allow GT years to be skipped in projections and use the projection framework to explore short- and long-term risks (local PDF, paragraph 54) (Commission for the Conservation of Southern Bluefin Tuna 2026).
2026 internal context
The subsequent 2027–2029 proposal resolves the immediate programme reference: GT is proposed annually, with about 5,000 releases and 13,000 harvest samples per cohort and a target CV of 0.25 (Preece and Eveson 2026). The 2025 and 2026 cohorts are already in the pipeline. The 2025 cohort has 3,680 releases and 17,000 collected harvest samples; the 2026 cohort has 3,066 releases, with its 13,000-sample harvest target due in 2027 (Preece et al. 2026; Preece and Eveson 2026). These observed and committed quantities replace the generic future sample sizes in the long-horizon comparison below.
Email discussion in April and May 2026 did not define a quantitative terms of reference. It identified two candidate forms of reduction–a smaller annual exercise and a lower survey frequency–and asked how either would affect the robustness of ESC advice. It also recommended separating this work from the communication of the new MCMC-based risk results and using 2026, if necessary, to make the code ready for the fuller MP review.
The 2025 paper’s then-current schedule included no tagging in 2026. That schedule should not be assumed here: the April 2026 correspondence states that the full 2026 exercise had been reinstated. Scenario years therefore remain to be specified from the latest agreed program rather than copied from the 2025 paper.
The July 2026 discussion with Ana adds two useful design cautions:
- exploratory results could help internal planning for next year even if they are not presented at ESC31; and
- past work found it difficult to show a large effect from GT gaps alone and obtained a stronger effect only when GT gaps were combined with failing CPUE, represented as a change in catchability q.
The first comparison should therefore isolate GT frequency under the standard CPUE assumptions. Any CPUE-q failure should be a clearly labelled, secondary joint-stress experiment, not part of the definition of the GT-only effect.
Define the calendar before defining scenarios
Three different years can be attached to one GT abundance estimate: Table 1 defines the terms used throughout this page.
| Term | Current data field | Meaning |
|---|---|---|
| Tagging or release year | RelYear |
Age-2 fish are released in this field year. |
| Harvest sampling or recapture year | RecYear |
The tagged cohort is sampled, currently represented as RelYear + 1. |
| Data-availability year | CTP schedule | The completed estimate becomes available to a later TAC calculation. |
Every scenario should be stored and reported in all three year conventions. This will prevent a nominal “2028 gap” from meaning a 2028 field season in one place and a 2028 data exchange in another.
The sbt projection interface now distinguishes gt_skip_years, which names release years, from gt_skip_recapture_years, which independently names harvest-sampling or recapture years. The generated template validates the release and recapture calendars and retains the one-year relationship RecYear = RelYear + 1 for rows that remain. Unit tests cover release-only, recapture-only, and combined omissions, including the fact that a literal shutdown of odd-calendar-year fieldwork can remove every otherwise retained even-year release under the current one-year lag.
Experiment A: consequences for the stock assessment
The first assessment experiment can use historical data omission as a controlled sensitivity analysis:
- retain the accepted ESC31 assessment as the baseline;
- remove GT rows according to agreed release-year patterns;
- refit the otherwise unchanged assessment;
- compare the posterior displacement and loss of precision; and
- project from each accepted posterior using the same future assumptions.
Suggested assessment summaries are:
- recruitment by cohort, especially cohorts whose GT estimate was withheld;
- spawning biomass, total reproductive output, depletion, and rebuilding probabilities;
- recruitment variability and other parameters that GT data may inform;
- changes in the posterior uncertainty carried into projections;
- GT likelihood and predictive diagnostics; and
- optimizer and MCMC diagnostics for every refit.
This retrospective exercise cannot estimate statistical bias because the true historical population is unknown. It measures sensitivity and loss of information. Bias, interval coverage, and repeated-assessment performance need simulated operating models with known truth.
Experiment B: consequences within projections and the CTP
For each accepted starting posterior draw, the projection experiment should:
- simulate the same underlying population and non-GT monitoring processes;
- apply an explicit release-year schedule to the simulated GT observations;
- pass the resulting missing estimates into the existing CTP update schedule;
- retain complete CTP diagnostics and TAC paths; and
- compare scenarios using paired posterior draws and, as far as possible, common random numbers.
Primary summaries should include:
- the GT estimates and numbers of matches actually available at every CTP calculation;
- the number and age of usable GT estimates in the CTP’s recent window;
- TAC level, TAC change, TAC variability, and cumulative catch;
- short- and long-term stock risks using the same definitions as the main projection report;
- the delay in management response to weak recruitment; and
- the frequency of CTP calculations with too little usable GT information.
The common-random-number safeguards are implemented. Projection simulations use component-specific random streams for CPUE, GT, HSP, and POP, and GT matches are keyed so that a retained release year receives the same draw whether or not other release years are omitted. Unit tests confirm that removing GT rows does not shift the HSP or POP draws. A paired comparison must still use the same posterior-row selection, component seeds, and non-GT inputs in both arms.
Programme-comparison scenarios
The four scenarios in Table 2 use the same accepted historical posterior. They differ only in future GT effort or availability. The annual full-effort scenario is the proposed programme; the other three are counterfactual precision and resilience tests.
| ID | GT design | Purpose |
|---|---|---|
| GT0 | Annual full effort: 5,000 releases and 13,000 harvest samples from 2027 | Proposed 2027–2029 programme and reference case. |
| GT1 | Annual reduced effort: 5,000 releases and 10,000 harvest samples from 2027 | Isolate the documented lower harvest-sample design while retaining annual estimates. |
| GT2 | Retain the 2027 release, then deliberately omit alternate release cohorts from 2028 | Counterfactual test of regular gaps; not an operational recommendation. |
| GT3 | GT2 plus unplanned failure of the 2029 release cohort | Stress test three consecutive missing cohort estimates (2028–2030); not a proposed programme. |
The formal proposal covers 2027–2029. For a controlled long-horizon comparison, each scenario’s 2027 design is held constant after 2029 through the end of the simulated programme. That continuation is an experimental assumption, not an approved fieldwork schedule.
The deliberate-skip arm omits release cohorts 2028, 2030, …, 2054 and their following-year recapture work. The combined stress arm makes those same deliberate omissions and adds an unplanned failure of the 2029 release, leaving the 2028–2030 cohort estimates consecutively unavailable.
All four use the standard CPUE process. A CPUE-q change is not mixed into this comparison; any joint-stress test would need a separately specified start year, size, duration, and biological rationale.
What the current projection code can do
Table 3 distinguishes implemented capabilities from the remaining extensions needed for a broader program-design analysis.
| Capability | Current status | Consequence for this experiment |
|---|---|---|
| Represent a missing future GT estimate in the CTP | Implemented, tested, and exercised in both the 100-draw screen and 2,000-draw long-horizon comparison below. Separate arguments omit release years and recapture years. A missing release-year estimate becomes gtN = NA and gtR = 0; the CTP weighted mean uses the remaining finite estimates with positive matches. |
The independent production review gate passed for all four programme scenarios. |
| Apply the OMMP16-style GT availability lag | Implemented and verified. The default CTP schedule uses a four-year lag from TAC implementation year to the latest GT release year. | The release, recapture, data-availability, and nine CTP implementation years were checked independently for every scenario. |
| Detect a window with no usable GT estimate | Implemented. The CTP stops rather than silently inventing an estimate. | Long or clustered gaps need deliberate failure-handling rules for a full MSE. |
| Apply one reduced release and harvest sample size to all future GT years | Implemented. | Supports a constant lower-effort sensitivity. |
| Apply different effort levels in different future years | Implemented and tested. Release and harvest sample-size vectors are keyed by release year. | The observed 2025–2026 pipeline is shared by all four long-horizon arms before the 2027 programme scenarios diverge. |
| Simulate a future CPUE-q failure or step change | Not exposed as a projection scenario. The current path carries forward fitted catchability and creep. | The optional joint-stress experiment needs a defined q trajectory and code support. |
| Refit the full stock assessment at future assessment years | Not implemented in run_projections(). It projects from a fixed accepted posterior and runs the CTP’s internal model, but does not update the full assessment posterior. |
Experiment A must use separate refits. A fully closed-loop assessment experiment would be a larger MSE extension. |
| Keep all non-GT stochastic draws paired when GT rows are removed | Implemented and tested. CPUE, GT, HSP, and POP use component-specific streams, and retained GT rows use stable release-year keys. | Use the same posterior rows, component seeds, and non-GT inputs in both arms. |
| Retain weak-stock draws when nominal TAC cannot be taken | Implemented, tested, and independently audited. The opt-in operating-model rule scales every within-year allocation proportionally to keep raw harvest at or below 0.9, while retaining nominal advice and realized catch separately. The package default remains fail-closed. | Report the probability and magnitude of catch shortfall; do not hide infeasible advice by dropping draws, and interpret stock risk alongside the lower realized removals. |
| Project far enough to see spawning-stock consequences | Implemented in the long-horizon comparison. Dynamics extend through 2055 and the population state through 2056. | This spans nine CTP updates and more than the documented 10–15+ year pathway from an age-2 cohort to spawning-stock outcomes. |
The code now contains the calendar, year-specific effort, common-posterior, and random-stream machinery for the four-arm long-horizon comparison. This is still an assessment-posterior-conditioned closed-loop CTP analysis, not a formal revalidation or retuning of the MP across the full CCSBT operating-model reference set.
Proposed stages
Stage 0: agree the question
- Decide whether the immediate question is CTP robustness, value to the stock assessment, a cost comparison, or preparation for a revised MP.
- Confirm the latest operational GT schedule and map release years to data availability.
- Agree whether results are internal screening material or formal ESC output.
Stage 1: make the projection comparison identifiable
Stage 2: small diagnostic screen
- Use this screen to decide whether to retain the scenario and refine the horizon and performance measures.
Stage 3: assessment sensitivity
Stage 4: production analysis or full MSE
- Add the CPUE-q joint stress only if it has a defensible specification.
- If the question becomes a change to the CTP or long-term program design, move from screening projections to a formal MSE and MP-review process.
Questions outside this comparison
- The analysis does not attach monetary costs to the four scenarios.
- A formal MP review must decide which operating-model reference set, low- recruitment stresses, performance measures, and tuning criteria are needed for full MSE revalidation.
- A CPUE-q joint stress remains separate unless evidence defines its magnitude, start year, and duration.
- Any proposed programme change still requires scientific review through the exceptional-circumstances and meta-rule process.
Working provenance
- OMMP16 report, projection workplan paragraph 54 (Commission for the Conservation of Southern Bluefin Tuna 2026).
- Preece, Davies, Galeano, Hillary, and Eveson (2025), especially sections 3–6 (Preece et al. 2025).
- Internal email thread, Exploring Gene Tagging Alternatives, 29 April–13 May 2026.
- Internal email thread, 2026 MP evaluation - need a breakdown to justify potential budget, August 2025.
- Darcy Webber–Ana Parma discussion, 23–24 July 2026.
The correspondence entries document planning context only. Formal scientific claims and any future terms of reference should be tied to approved CCSBT documents.
Earlier historical-omission comparison
The earlier run retains the accepted ESC31 base fit as the reference and changes GT availability in both the alternative historical assessment and its future monitoring schedule. The base MLE was not refitted. The first gate was an alternative MLE using the same model configuration, priors, parameter map, optimizer, and biological and numerical acceptance criteria. That MLE, its posterior, and the paired diagnostic projection are now complete.
The posterior review below uses all 3,000 retained draws from each fit. The initial projection comparison uses 100 chain-balanced posterior draws from each fit. It is retained as a diagnostic of historical information loss, not as the programme recommendation or the main future-monitoring comparison. Its different historical posteriors are specifically avoided in the common- posterior long-horizon analysis below.
Selected alternate-cohort design
The historical-omission experiment has one implementation: odd_release_cohorts. It drops the complete GT estimate for releases in 2017, 2019, 2021, and 2023, including the associated recapture work one year later. It retains the four GT estimates from releases in 2016, 2018, 2022, and 2024.
This is the requested every-second-estimate comparison. Its operational calendar is important: retained even-year releases are recaptured in the following odd calendar year. It therefore does not represent a complete shutdown of all fieldwork in odd calendar years. A literal odd-calendar-year shutdown would remove the recaptures for every retained even-year release and, with the current one-year release–recapture lag, leave no usable GT estimates.
Reproducible run contract
- Fit artifacts are written under
ESC31/runs/gtskip/; failed MCMC attempts and their diagnostics are retained rather than overwritten invisibly. - The first review gate stops after the alternative MLE; no MCMC or projection is launched until that fit is accepted for further work.
- The alternative MLE must have convergence code zero, maximum gradient no greater than 0.01, estimability, valid biological states, exact catch accounting, and the accepted LL4 configuration.
- The alternative MCMC uses four chains, 150 warmup and 750 retained iterations per chain, dense metric, the MLE mode as the chain start,
adapt_delta = 0.999, maximum treedepth 13, and seed 73015. - Posterior acceptance requires maximum rank-normalized R-hat below 1.01, bulk and tail ESS at least 400, no divergences, no maximum-treedepth hits, and valid biological states for every retained draw.
- The base and alternative 100-draw projections use the same retained iteration numbers in each chain and common component-specific random streams. CPUE, GT, HSP, and POP streams are isolated so removing GT years cannot shift the HSP or POP random draws.
- Projected skip controls distinguish release years from harvest-sampling or recapture years. The generated calendar and the GT rows reaching each CTP update will be saved and reported.
- The fit, MCMC, and paired projection artifacts are retained separately from the accepted production assessment and projection caches.
Only files carrying the odd_release_cohorts experiment name are inputs or outputs of the historical-omission analysis documented here. The older unprefixed gtskip_mle_* files and esc31_gtskip_odd_calendar_fieldwork.sbt.rds are preserved exploratory artifacts for different designs; they are not accepted results and must not be substituted for the named files. The local artifact registry in runs/gtskip/README.md records that distinction.
MLE review
The alternate-release-cohort alternative has been fitted to maximum likelihood and passes every numerical, estimability, biological-state, catch-accounting, harvest-wall, and LL4 gate. Its posterior and paired diagnostic projection are reviewed separately below.
The executable comparison retains four of the eight historical GT observations. An independent input audit confirms that the numerical contents of all other model data, the parameter map, priors, bounds, and optimizer controls match the accepted base. Table 4 reports the acceptance checks. The alternative was initialized from the base MLE, but every active parameter was then re-estimated.
| Check | Requirement | Base | Alternate-cohort GT MLE |
|---|---|---|---|
| Historical GT observations | 8 versus 4 | 8 | 4 |
| Final negative log-likelihood | Finite | 7616.637822 | 7600.761697 |
| Optimizer convergence code | 0 | 0 | 0 |
| Maximum absolute gradient | ≤ 0.01 | 2.60 × 10⁻¹⁰ | 5.72 × 10⁻⁵ |
| Estimability | Pass | Pass | Pass |
| Biological-state gate | Pass | Pass | Pass |
| Maximum raw seasonal harvest at age | ≤ 0.9 | 0.617533 | 0.609524 |
| Minimum populated numbers-at-age | > 0 | 1696.516 | 1723.101 |
| Maximum scaled catch error | ≤ 10⁻⁸ | 2.53 × 10⁻¹⁶ | 3.65 × 10⁻¹⁶ |
| Harvest-wall objective contribution | Report (not a gate) | 2.57 × 10⁻²¹ | 5.17 × 10⁻²² |
| Continuation contribution | ≤ 10⁻⁸ | 0 | 0 |
| LL4 standard-fleet contract | Pass | Pass | Pass |
The alternative objective is 15.876 units lower, but this is not evidence of a better-fitting model: its objective omits the four GT likelihood contributions from the omitted release cohorts. The useful comparison is the change in fitted stock quantities in Table 5, not the raw objective difference.
| MLE quantity | Base | Alternate-cohort GT | Difference | Relative difference |
|---|---|---|---|---|
| B₀ | 6,422,189 | 6,426,113 | +3,924 | +0.06% |
| 2025 TRO | 1,653,497 | 1,777,381 | +123,884 | +7.49% |
| 2025 relative TRO | 0.2575 | 0.2766 | +0.0191 | +7.43% |
| 2025 recruitment | 3,216,124 | 3,405,529 | +189,405 | +5.89% |
| 2026 TRO | 1,717,909 | 1,860,853 | +142,944 | +8.32% |
| 2026 relative TRO | 0.2675 | 0.2896 | +0.0221 | +8.25% |
| 2026 recruitment | 3,645,944 | 3,878,668 | +232,723 | +6.38% |
| M₀ | 0.36574 | 0.36969 | +0.00395 | +1.08% |
| M₄ | 0.16167 | 0.16342 | +0.00175 | +1.08% |
| M₁₀ | 0.10736 | 0.10852 | +0.00116 | +1.08% |
| M₃₀ | 0.45771 | 0.45994 | +0.00223 | +0.49% |
| CPUE q | 0.95749 | 0.96826 | +0.01078 | +1.13% |
Retaining every second GT estimate still raises the recent fitted recruitment peaks and the terminal stock trajectory in this MLE (Figure 1). The largest proportional recruitment change over the fitted trajectory is about 38.9% in 2015; by 2026, recruitment is 6.4% higher and relative TRO is 0.0221 higher. The posterior review below evaluates uncertainty around the assessment trajectories, and the 100-draw screen then evaluates the first CTP response. The MLE result alone is not a management conclusion. The accepted ESC31 base, grid, and production projections remain closed while this separate robustness analysis continues.
The accepted fit is runs/gtskip/esc31_gtskip_odd_release_cohorts.sbt.rds. Reproducible review tables and the plotted trajectories are stored beside it as gtskip_odd_release_cohorts_mle_diagnostics.csv, gtskip_odd_release_cohorts_mle_management.csv, and gtskip_odd_release_cohorts_mle_review.rds.
Alternate-cohort posterior review
The local four-chain alternate-cohort MCMC passed every numerical and biological-state acceptance gate. It is an accepted assessment sensitivity posterior for this screen, not a GT-skipping projection or management procedure result.
The sampler used the pre-specified dense-metric contract: four chains, 150 warmup iterations and 750 retained iterations per chain, exact MLE-mode starts, adapt_delta = 0.999, maximum treedepth 13, and seed 73015. Table 6 records the resulting diagnostics. The initial automatic step-size search had a few warmup-only divergences; the 3,000 retained draws used for inference had none.
| Check | Requirement | Alternate-cohort posterior |
|---|---|---|
| Chains | 4 | 4 |
| Warmup per chain | 150 | 150 |
| Retained per chain | 750 | 750 |
| Total retained draws | 3,000 | 3,000 |
| Maximum rank-normalized R-hat | < 1.01 | 1.0076 |
| Minimum bulk ESS | ≥ 400 | 1,632.1 |
| Minimum tail ESS | ≥ 400 | 1,153.3 |
| Divergences after warmup | 0 | 0 |
| Maximum-treedepth hits | 0 | 0 |
| State draws expected / checked | 3,000 / 3,000 | 3,000 / 3,000 |
| Invalid or non-finite state draws | 0 | 0 |
| Maximum raw seasonal harvest at age | ≤ 0.9 | 0.8731 |
| Minimum populated numbers-at-age | > 0 | 603.348 |
| Maximum continuation penalty | ≤ 10⁻⁸ | 0 |
| Overall posterior gate | Pass | Pass |
Historical GT rows
The alternate fit skipped releases in 2017, 2019, 2021, and 2023 and their associated recapture years 2018, 2020, 2022, and 2024. It retained releases in 2016, 2018, 2022, and 2024, recaptured respectively in 2017, 2019, 2023, and 2025. This is the linked cohort-level schedule in Table 7; it is not a blanket shutdown of all work in odd calendar years.
There is no 2020 release row in Table 7 because tagging was cancelled in that field year. That source-data gap is separate from the retained/omitted pattern imposed by this sensitivity.
| Release year | Release age | Recapture year | Releases | Scanned samples | Matches | Alternate-fit status |
|---|---|---|---|---|---|---|
| 2016 | 2 | 2017 | 2,952 | 15,389 | 20 | Retained |
| 2017 | 2 | 2018 | 6,480 | 11,932 | 67 | Omitted |
| 2018 | 2 | 2019 | 6,295 | 11,980 | 66 | Retained |
| 2019 | 2 | 2020 | 4,242 | 11,109 | 31 | Omitted |
| 2021 | 2 | 2022 | 6,401 | 10,742 | 41 | Omitted |
| 2022 | 2 | 2023 | 5,084 | 14,714 | 38 | Retained |
| 2023 | 2 | 2024 | 2,759 | 13,297 | 14 | Omitted |
| 2024 | 2 | 2025 | 3,522 | 11,011 | 11 | Retained |
Posterior trajectories and GT fits
Figure 2 compares the assessment trajectories using all 3,000 retained draws from each posterior. These are separately sampled base and alternate assessment distributions; they are not paired projection draws.
At 2026, median relative TRO is 0.3044 in the alternate fit versus 0.2817 in the base, an absolute difference of 0.0227 (8.0% relative to the base median). The corresponding 95% intervals are 0.2447–0.3829 and 0.2216–0.3526. Median recruitment is 3.819 million versus 3.680 million, a difference of 0.139 million (3.8%); its two wide 95% intervals also overlap substantially. The upward terminal-TRO direction seen at the MLE therefore remains in the posterior medians, but the assessment uncertainty is material and this is not a paired test of projected outcomes.
The fitted GT expectations in Figure 3 use the sample sizes in Table 7. Orange points are observations used in the corresponding likelihood. Grey crosses locate the four observations omitted from the alternate fit; they are shown only for context and did not contribute to its likelihood.
Paired 100-draw CTP projection screen
The annual-GT reference and alternate-release-cohort arms each completed 100 chain-balanced draws. All 200 historical states and all 400 CTP update optimisations across the two arms passed their configured gates. This accepts the run as a diagnostic screen; it does not promote it to the 2,000-draw production assessment or make it ESC advice.
The comparison changes two linked parts of the workflow: the alternate arm starts from the accepted historical-omission posterior reported above and also omits the corresponding future GT release cohorts. It therefore represents a continued alternate-cohort program, not a projection-only test in which both arms start from the same posterior.
The 100 rows comprise 25 retained iterations from each of four chains. Both arms use draw seed 44, projection seed 102, the same retained iteration numbers, the same recruitment and selectivity rules, the same fixed TAC and fleet allocation, and component-specific common random streams. The projection covers model and monitoring years 2022–2035 and the start-of-2036 population state.
Future GT calendar
The future schedule in Table 8 skips releases in 2025, 2027, 2029, 2031, and 2033 and their linked recaptures in 2026, 2028, 2030, 2032, and 2034. Retained releases still require recapture work in the following odd calendar year.
This retained 100-draw screen is deliberately counterfactual from release year 2025: it skips the 2025 release even though that field programme has already occurred. It is therefore not the current operational pipeline. It also omits odd release years, whereas the later common-posterior GT2 and GT3 arms retain 2027 and omit even release years from 2028. “Alternate cohort” describes the frequency in both experiments, not a shared parity or calendar.
| Release year | Recapture year | Alternate-arm GT estimate |
|---|---|---|
| 2025 | 2026 | Skipped |
| 2026 | 2027 | Retained |
| 2027 | 2028 | Skipped |
| 2028 | 2029 | Retained |
| 2029 | 2030 | Skipped |
| 2030 | 2031 | Retained |
| 2031 | 2032 | Skipped |
| 2032 | 2033 | Retained |
| 2033 | 2034 | Skipped |
| 2034 | 2035 | Retained |
Projection diagnostics
Table 9 records the complete numerical gate. The alternate arm’s maximum CTP gradient, (4.9811^{-4}), passes the (5^{-4}) tolerance but is close enough that it should be watched in a larger run. Every final CTP Hessian was positive definite and every historical state passed.
| Check | Requirement | Annual GT | Alternate release cohorts |
|---|---|---|---|
| Projection draws expected | 100 | 100 | 100 |
| Historical states checked | 100 | 100 | 100 |
| Historical states passed | 100 | 100 | 100 |
| CTP updates expected | 200 | 200 | 200 |
| CTP updates passed | 200 | 200 | 200 |
| Positive-definite CTP Hessians | 200 | 200 | 200 |
| Maximum absolute CTP gradient | ≤ 0.0005 | 0.0004668 | 0.0004981 |
| Elapsed time | Report only | 144.8 s | 144.9 s |
CTP inputs and TAC response
The first CTP update in 2030 can use GT information through release year 2026; the second, in 2033, can use information through 2029. Table 10 shows that skipping estimates did not simply lower the weighted GT abundance. The remaining match-weighted estimates produced a slightly higher median at both updates, and the median GT multiplier stayed at its neutral value of one in both arms.
| CTP implementation year | Latest GT release year | Median GT abundance: annual | Median GT abundance: alternate | Median GT multiplier: annual / alternate | Median TAC: annual | Median TAC: alternate | Decrease / unchanged / increase: annual | Decrease / unchanged / increase: alternate |
|---|---|---|---|---|---|---|---|---|
| 2030 | 2026 | 2,244,647 | 2,288,038 | 1.00 / 1.00 | 26,647.0 | 26,647.0 | 0 / 0 / 100 | 0 / 0 / 100 |
| 2033 | 2029 | 1,953,125 | 2,222,222 | 1.00 / 1.00 | 26,802.3 | 26,818.8 | 1 / 37 / 62 | 2 / 30 / 68 |
Figure Figure 4 shows the resulting marginal TAC paths.
The central TAC paths are very similar, but the result is more nuanced than “no effect”:
- The 2026–2029 TACs are prescribed and therefore identical.
- In 2030, the median unconstrained TAC is about 35,542 t in the annual arm and 35,502 t in the alternate arm. Every draw reaches the 3,000 t increase limit, so both implemented medians are exactly 26,647 t through 2032.
- At the 2033 update, the marginal medians differ by only 16.5 t (26,802.3 versus 26,818.8 t). Their empirical 95% intervals are 26,647–29,647 t and 26,393–29,647 t, respectively.
- The draw-level responses are not identical: 74 of 100 paired computational rows differ at the 2033 update, and the empirical paired difference interval is -2,852 to +2,540 t. This artificial common-random-number pairing is useful for a screen, but it is not a credible interval from a formally joint posterior.
- Median cumulative nominal TAC over 2026–2035 is 253,960 t for annual GT and 253,998 t for the alternate arm. The paired median difference is zero, while its empirical 95% interval is -8,557 to +7,619 t.
The first update is therefore uninformative about the operational effect of GT gaps because the TAC-change cap binds. The second update suggests very similar central TAC advice but enough draw-level variation that equivalence should not be claimed from 100 draws.
Projected stock status and risk
Figure Figure 5 and Table Table 11 report the corresponding stock-status screen.
| State year | Annual GT median (95% interval) | Alternate median (95% interval) | Difference in marginal medians |
|---|---|---|---|
| 2030 | 0.330 (0.265–0.421) | 0.363 (0.280–0.461) | +0.034 |
| 2033 | 0.345 (0.262–0.498) | 0.387 (0.289–0.516) | +0.042 |
| 2036 | 0.344 (0.241–0.571) | 0.384 (0.270–0.546) | +0.040 |
The alternate arm retains the higher terminal stock-status direction already seen in its historical assessment posterior. No draw in either arm falls below relative TRO 0.2 during the configured short-term window (2030–2032) or longer window (2033–2036). With 100 draws this means only that the observed screening frequency was zero; it does not establish a zero risk probability.
The higher alternate trajectory must not be attributed solely to future GT skipping. It also reflects conditioning on the different accepted historical posterior, and the intervals overlap substantially. Moreover, the current horizon is too short to capture the full documented 10–15+ year path from a future age-2 cohort to spawning-stock outcomes.
Interpretation and next gate
The screen supports three provisional comments:
- Removing alternate GT estimates did not automatically depress the CTP’s weighted GT input; which cohorts remain and their match weights matter.
- The median TAC path was effectively unchanged through the first update and differed by only 16.5 t at the second, largely because the first update hit the TAC-change cap and the median GT multiplier remained neutral.
- Individual draw responses and cumulative differences were much wider than the marginal medians suggest, while the stock-status comparison also includes historical-posterior displacement.
Those limitations motivate the common-posterior long-horizon comparison below. The earlier result remains useful for showing that historical GT omission can move the assessment posterior, but its 16.5 t median TAC difference is not used to judge whether a future programme is acceptable.
Common-posterior long-horizon closed-loop comparison
This comparison isolates future monitoring. Every arm uses the same 2,000 chain-balanced draws from the accepted base posterior, the same future recruitment and selectivity paths, the same catch allocation and removal assumptions, and component-specific common random numbers. Only the future GT effort and availability schedule changes.
The projection covers dynamics from 2022 through 2055 and the resulting state through 2056. Fixed TACs apply from 2026 through 2029, after which the CTP is run at nine implementation years: 2030, 2033, 2036, 2039, 2042, 2045, 2048, 2051, and 2054. This removes the two-update horizon limitation of the earlier screen and spans the 10–15+ year pathway from juvenile recruitment to the spawning stock.
Nominal CTP TAC advice and realized catch are retained separately. No posterior draw is discarded or replaced if its stock becomes too small to take the nominal TAC. Instead, an explicit operating-model implementation rule reduces all seasonal and fishery allocations in that draw-year by one common factor until the raw combined seasonal harvest rate is no greater than 0.9. The report therefore audits catch shortfall as an outcome. Because this rule can reduce late-horizon removals in weak-stock draws, risk results must be interpreted together with both nominal TAC and realized catch rather than as an unqualified consequence of the monitoring design.
The programme’s target CV of 0.25 is retained as an operational design target, not imposed as a model acceptance gate. The results report simulated match counts and use 16 matches (the crude Poisson quantity (1/0.25^2)) only as a descriptive precision proxy. That proxy is not a replacement for a formal design-based CV calculation.
This is a closed-loop, posterior-conditioned CTP simulation: monitoring data are generated from each operating-model path, passed to the CTP, and the resulting TAC changes feed back into subsequent population and observation paths. It is substantially more informative than the 100-draw screen. It is not, however, formal revalidation or retuning of the MP across the complete CCSBT operating-model reference set, and it cannot authorize a change from the annual programme.
All four 2,000-draw arms passed the independent numerical and biological review. Reducing the annual harvest sample lowered simulated match precision but retained five recent estimates in every CTP window and produced results very close to the full-effort reference in this experiment. Deliberate gaps reduced the recent-information window to two or three estimates, and the added failure reduced it to one estimate at the 2036 CTP calculation. Those gaps increased TAC variability and gave slightly poorer long-horizon stock-risk indicators, although the paired simulation intervals were broad and included zero. These counterfactuals do not provide a basis for reducing the proposed annual programme.
Independent validation
The production run reused the same 2,000 chain-balanced posterior rows and paired non-GT stochastic inputs in all arms. Each arm passed all 2,000 historical-state checks and all 18,000 CTP optimisations. Every final CTP Hessian was positive definite, the maximum absolute gradient met the pre-specified \(5 \times 10^{-4}\) tolerance, populated abundance remained positive, raw harvest reached but did not exceed the configured 0.9 ceiling beyond its \(10^{-10}\) numerical tolerance, and the feasibility penalty was effectively zero. Table Table 12 records these gates.
| Programme scenario | Historical states passed | CTP updates passed | Maximum CTP gradient | Minimum Hessian eigenvalue | Overall gate |
|---|---|---|---|---|---|
| Annual full effort | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499976 | 12.025 | Pass |
| Annual reduced effort | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499976 | 11.897 | Pass |
| Deliberate alternate-cohort skips | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499930 | 12.051 | Pass |
| Deliberate skips plus unplanned failure | 2,000 / 2,000 | 18,000 / 18,000 | 0.000499953 | 12.034 | Pass |
Monitoring coverage and simulated precision
Figure Figure 6 and Table Table 13 separate the number of usable cohort estimates from the precision of those retained.
Reducing annual harvest sampling from 13,000 to 10,000 lowered the median simulated matches per active cohort from 29 to 22 and more than doubled the mean frequency below the descriptive 16-match precision proxy. It did not remove an estimate from the CTP window. The skip scenarios retained the full-effort precision for cohorts that were actually sampled, but had far fewer cohort estimates available. Match precision for an observed cohort and coverage of recent cohorts are therefore distinct consequences.
| Programme scenario | Median matches per active cohort | Probability below 16-match proxy | Minimum median usable estimates per CTP window |
|---|---|---|---|
| Annual full effort | 29 | 10.9% | 5 |
| Annual reduced effort | 22 | 23.2% | 5 |
| Deliberate alternate-cohort skips | 29 | 10.9% | 2 |
| Deliberate skips plus unplanned failure | 29 | 10.9% | 1 |
TAC response
Figures Figure 7 and Figure 8 show marginal and paired TAC responses; Table Table 14 summarizes the paired cumulative differences.
All arms are identical through the capped 2030 response. Later marginal median TACs are higher under deliberate gaps, especially after the added failure, but this is not evidence that missing monitoring improves performance. Missing cohorts change the recent match-weighted GT input in either direction rather than imposing a downward adjustment. The median paired cumulative increase is 1,543 t with deliberate skips and 3,332 t after the added failure, only about 0.2% and 0.4%, respectively, of the annual-full median cumulative TAC of 819,398 t. Their paired 95% intervals are much wider and span zero. The higher central nominal advice is also accompanied by slightly lower central stock status and higher long-term risk.
| Programme scenario | Median paired cumulative TAC difference (t) | Empirical paired 95% interval (t) | Mean absolute paired difference (t) | Draws with a different cumulative TAC |
|---|---|---|---|---|
| Annual reduced effort | 0 | -16,116 to 16,911 | 4,059 | 73.8% |
| Deliberate alternate-cohort skips | +1,543 | -43,030 to 52,483 | 14,348 | 86.0% |
| Deliberate skips plus unplanned failure | +3,332 | -58,170 to 72,301 | 20,413 | 88.2% |
Stock status and risk
Figure Figure 9 and Table Table 15 show the long-horizon stock distributions and the associated risk summaries.
The annual reduced-effort arm remains very close to full annual effort. The two gap scenarios have progressively lower 2056 median relative TRO and higher full-horizon probabilities of ever falling below 0.2. The central changes are small relative to the simulation uncertainty: even the combined skip-plus- failure arm has a paired 2056 difference interval of -0.0672 to +0.0588. Separation appears mainly in the final years, consistent with the delay from age-2 recruitment information to spawning-stock consequences.
| Programme scenario | 2056 relative TRO, median (95% interval) | Paired median difference from annual full (95% interval) | P(2056 relative TRO < 0.2) | P(ever < 0.2), 2030–2056 |
|---|---|---|---|---|
| Annual full effort | 0.291 (0.046–0.637) | Reference | 23.9% | 26.6% |
| Annual reduced effort | 0.290 (0.047–0.640) | 0.0000 (-0.0154 to +0.0146) | 23.8% | 26.7% |
| Deliberate alternate-cohort skips | 0.287 (0.047–0.636) | -0.0011 (-0.0478 to +0.0400) | 25.3% | 28.4% |
| Deliberate skips plus unplanned failure | 0.283 (0.043–0.638) | -0.0026 (-0.0672 to +0.0588) | 26.3% | 29.4% |
In the first decade (2030–2039), the probability of ever falling below 0.2 is about 2.8% in every arm. In the final 2050–2056 window it rises from 25.6% under annual full effort to 27.4% with deliberate skips and 28.6% with the additional failure. Across the complete 2030–2056 horizon, the corresponding ever-below risk rises from 26.6% to 28.4% (+1.9 percentage points) and 29.4% (+2.9 percentage points), using the underlying unrounded probabilities (1.85 and 2.85 percentage points before rounding). The probability of being below 0.2 specifically in 2056 rises by 1.4 and 2.4 percentage points. This analysis therefore detects a modest adverse long-term direction from missing monitoring, not a sharply separated short-term effect; the paired final-stock intervals still span zero.
Nominal advice and realized catch
Figure Figure 10 shows where nominal TAC advice cannot be fully realized under the operating-model harvest ceiling.
Median cumulative catch shortfall is zero in every arm, and only 2.5–2.6% of draws experience any shortfall. The 97.5th percentile of cumulative shortfall is about 7 t under either annual scenario, 389 t under deliberate skips, and 1,215 t after the added failure. A few weak-stock paths have much larger annual shortfalls. Stock-risk results must therefore be interpreted alongside realized catch: the retained weak-stock draws are not allowed to remove an infeasible nominal TAC.
Scientific interpretation
- Lower annual sample size primarily affects precision. It retains five recent GT estimates in each CTP window but roughly doubles the frequency of active cohorts falling below the descriptive match-count proxy. In this isolated GT-only comparison, its TAC and stock-risk results remain close to annual full effort.
- Skipping cohorts primarily affects coverage. Regular gaps reduce the five-cohort window to two or three estimates. An adjacent unplanned failure briefly leaves only one, increasing TAC dispersion and producing the largest adverse central stock-risk changes.
- The direction is more informative than the apparent separation. The gap scenarios have higher central nominal TACs and slightly lower relative TRO, but paired intervals span zero and the long-horizon stock distributions overlap strongly. Neither equivalence nor a precise causal effect should be claimed from this one posterior-conditioned experiment.
- Annual GT remains the appropriate reference. The comparison excludes a CPUE-q failure, future full-assessment refits, alternative operating-model reference sets, and MP retuning. Any proposed reduction still requires the formal MSE, MP-review, exceptional-circumstances, and meta-rule process.
The independent summaries and figures were generated with scripts/summarise-esc31-gtprogramme-mse.R from the atomic four-arm comparison and its four scenario caches. The review binds the exact input artifacts by MD5 and stops if calendars, paired non-GT inputs, biological states, CTP optimisations, or catch accounting fail their configured gates.