ESC31 Gene-tagging Programme Resilience Analysis

ImportantContinue annual gene tagging wherever possible

GT skipping is not an operational recommendation from this analysis. The 2027–2029 proposal specifies an annual programme of approximately 5,000 age-2 releases, 13,000 age-3 harvest samples, and a target CV of 0.25 (Preece and Eveson 2026). That annual full-effort programme is the reference case on this page. Deliberate gaps and an unplanned failure are counterfactual resilience tests: they show what could happen if monitoring is reduced or disrupted, not that alternate-year GT is acceptable.

Purpose

The analysis examines the value and resilience of the annual GT programme and the scientific and management consequences of reduced effort or missing field years. Two consequences need to be kept separate:

  1. Stock assessment: less information about individual recruiting cohorts may alter estimated recruitment, stock status, uncertainty, and the assessment posterior used to condition future operating models.
  2. Projection and management procedure: missing recent recruitment estimates may alter the GT component of the Cape Town Procedure (CTP), TAC paths, and short- and long-term risk.

The objective is to compare monitoring programmes transparently while keeping the accepted assessment and the recommended annual programme unchanged. Any proposal to alter GT frequency would remain an exceptional circumstance and would require the formal MSE, MP-review, and meta-rule process.

Evidence and previous work

Formal work already completed

The 2025 review of changes to the GT program (official PDF) considered three distinct cost-saving mechanisms (Preece et al. 2025):

  • less frequent tagging, specifically every second year or two years out of three;
  • smaller release and harvest sample sizes, with lower precision; and
  • fewer fieldwork days, with a higher chance of failing to obtain enough releases.

These are not interchangeable experiments. Skipping a complete year creates a missing abundance estimate; reducing effort usually retains an estimate but makes it less precise and increases the risk of complete failure.

The same review explains why recruitment information matters. GT is currently the only fishery-independent absolute abundance estimate for age-2 fish used in the assessment and CTP. Without an informative recruitment index, a poor cohort may not be clearly detected until it affects the spawning stock roughly 10–15 or more years later. The earlier aerial-survey work is also relevant: MSE testing following its cancellation indicated delayed rebuilding and lower average catch without a recruitment index (Preece et al. 2025).

For the current CTP, the recent GT abundance estimates are combined using numbers of matches as weights. A skipped year therefore removes information from the recent window rather than being replaced with an average observation. The 2025 review anticipated greater stochastic variation in TAC advice when there are gaps, and noted that maintaining the same risk criteria during MP retuning would generally reduce mean TAC.

There is also one useful observed precedent. Cancellation of tagging in 2020 left the 2018 birth cohort–which would have been released at age 2 in 2020–without a GT abundance estimate. The CTP was designed to continue with an occasional missing GT value, but the 2023 assessment discussion noted that the estimated strength of this cohort was instead informed mainly by the high 2022 CPUE value and catches of ages 3 and 4. This does not demonstrate a general GT-by-CPUE effect, but it explains why a separate CPUE-q interaction is scientifically relevant.

OMMP16 subsequently placed the following item directly in the 2026 projection workplan: allow GT years to be skipped in projections and use the projection framework to explore short- and long-term risks (local PDF, paragraph 54) (Commission for the Conservation of Southern Bluefin Tuna 2026).

2026 internal context

The subsequent 2027–2029 proposal resolves the immediate programme reference: GT is proposed annually, with about 5,000 releases and 13,000 harvest samples per cohort and a target CV of 0.25 (Preece and Eveson 2026). The 2025 and 2026 cohorts are already in the pipeline. The 2025 cohort has 3,680 releases and 17,000 collected harvest samples; the 2026 cohort has 3,066 releases, with its 13,000-sample harvest target due in 2027 (Preece et al. 2026; Preece and Eveson 2026). These observed and committed quantities replace the generic future sample sizes in the long-horizon comparison below.

Email discussion in April and May 2026 did not define a quantitative terms of reference. It identified two candidate forms of reduction–a smaller annual exercise and a lower survey frequency–and asked how either would affect the robustness of ESC advice. It also recommended separating this work from the communication of the new MCMC-based risk results and using 2026, if necessary, to make the code ready for the fuller MP review.

The 2025 paper’s then-current schedule included no tagging in 2026. That schedule should not be assumed here: the April 2026 correspondence states that the full 2026 exercise had been reinstated. Scenario years therefore remain to be specified from the latest agreed program rather than copied from the 2025 paper.

The July 2026 discussion with Ana adds two useful design cautions:

  • exploratory results could help internal planning for next year even if they are not presented at ESC31; and
  • past work found it difficult to show a large effect from GT gaps alone and obtained a stronger effect only when GT gaps were combined with failing CPUE, represented as a change in catchability q.

The first comparison should therefore isolate GT frequency under the standard CPUE assumptions. Any CPUE-q failure should be a clearly labelled, secondary joint-stress experiment, not part of the definition of the GT-only effect.

Define the calendar before defining scenarios

Three different years can be attached to one GT abundance estimate: Table 1 defines the terms used throughout this page.

Table 1: Calendar terms used for gene-tagging release, recapture, and data availability.
Term Current data field Meaning
Tagging or release year RelYear Age-2 fish are released in this field year.
Harvest sampling or recapture year RecYear The tagged cohort is sampled, currently represented as RelYear + 1.
Data-availability year CTP schedule The completed estimate becomes available to a later TAC calculation.

Every scenario should be stored and reported in all three year conventions. This will prevent a nominal “2028 gap” from meaning a 2028 field season in one place and a 2028 data exchange in another.

ImportantProjection calendar correction complete

The sbt projection interface now distinguishes gt_skip_years, which names release years, from gt_skip_recapture_years, which independently names harvest-sampling or recapture years. The generated template validates the release and recapture calendars and retains the one-year relationship RecYear = RelYear + 1 for rows that remain. Unit tests cover release-only, recapture-only, and combined omissions, including the fact that a literal shutdown of odd-calendar-year fieldwork can remove every otherwise retained even-year release under the current one-year lag.

Experiment A: consequences for the stock assessment

The first assessment experiment can use historical data omission as a controlled sensitivity analysis:

  1. retain the accepted ESC31 assessment as the baseline;
  2. remove GT rows according to agreed release-year patterns;
  3. refit the otherwise unchanged assessment;
  4. compare the posterior displacement and loss of precision; and
  5. project from each accepted posterior using the same future assumptions.

Suggested assessment summaries are:

  • recruitment by cohort, especially cohorts whose GT estimate was withheld;
  • spawning biomass, total reproductive output, depletion, and rebuilding probabilities;
  • recruitment variability and other parameters that GT data may inform;
  • changes in the posterior uncertainty carried into projections;
  • GT likelihood and predictive diagnostics; and
  • optimizer and MCMC diagnostics for every refit.

This retrospective exercise cannot estimate statistical bias because the true historical population is unknown. It measures sensitivity and loss of information. Bias, interval coverage, and repeated-assessment performance need simulated operating models with known truth.

Experiment B: consequences within projections and the CTP

For each accepted starting posterior draw, the projection experiment should:

  1. simulate the same underlying population and non-GT monitoring processes;
  2. apply an explicit release-year schedule to the simulated GT observations;
  3. pass the resulting missing estimates into the existing CTP update schedule;
  4. retain complete CTP diagnostics and TAC paths; and
  5. compare scenarios using paired posterior draws and, as far as possible, common random numbers.

Primary summaries should include:

  • the GT estimates and numbers of matches actually available at every CTP calculation;
  • the number and age of usable GT estimates in the CTP’s recent window;
  • TAC level, TAC change, TAC variability, and cumulative catch;
  • short- and long-term stock risks using the same definitions as the main projection report;
  • the delay in management response to weak recruitment; and
  • the frequency of CTP calculations with too little usable GT information.

The common-random-number safeguards are implemented. Projection simulations use component-specific random streams for CPUE, GT, HSP, and POP, and GT matches are keyed so that a retained release year receives the same draw whether or not other release years are omitted. Unit tests confirm that removing GT rows does not shift the HSP or POP draws. A paired comparison must still use the same posterior-row selection, component seeds, and non-GT inputs in both arms.

Programme-comparison scenarios

The four scenarios in Table 2 use the same accepted historical posterior. They differ only in future GT effort or availability. The annual full-effort scenario is the proposed programme; the other three are counterfactual precision and resilience tests.

Table 2: Gene-tagging scenarios in the common-posterior long-horizon comparison. Only GT0 describes the proposed programme.
ID GT design Purpose
GT0 Annual full effort: 5,000 releases and 13,000 harvest samples from 2027 Proposed 2027–2029 programme and reference case.
GT1 Annual reduced effort: 5,000 releases and 10,000 harvest samples from 2027 Isolate the documented lower harvest-sample design while retaining annual estimates.
GT2 Retain the 2027 release, then deliberately omit alternate release cohorts from 2028 Counterfactual test of regular gaps; not an operational recommendation.
GT3 GT2 plus unplanned failure of the 2029 release cohort Stress test three consecutive missing cohort estimates (2028–2030); not a proposed programme.

The formal proposal covers 2027–2029. For a controlled long-horizon comparison, each scenario’s 2027 design is held constant after 2029 through the end of the simulated programme. That continuation is an experimental assumption, not an approved fieldwork schedule.

The deliberate-skip arm omits release cohorts 2028, 2030, …, 2054 and their following-year recapture work. The combined stress arm makes those same deliberate omissions and adds an unplanned failure of the 2029 release, leaving the 2028–2030 cohort estimates consecutively unavailable.

All four use the standard CPUE process. A CPUE-q change is not mixed into this comparison; any joint-stress test would need a separately specified start year, size, duration, and biological rationale.

What the current projection code can do

Table 3 distinguishes implemented capabilities from the remaining extensions needed for a broader program-design analysis.

Table 3: Current implementation capability for assessment and projection gene-tagging skip experiments.
Capability Current status Consequence for this experiment
Represent a missing future GT estimate in the CTP Implemented, tested, and exercised in both the 100-draw screen and 2,000-draw long-horizon comparison below. Separate arguments omit release years and recapture years. A missing release-year estimate becomes gtN = NA and gtR = 0; the CTP weighted mean uses the remaining finite estimates with positive matches. The independent production review gate passed for all four programme scenarios.
Apply the OMMP16-style GT availability lag Implemented and verified. The default CTP schedule uses a four-year lag from TAC implementation year to the latest GT release year. The release, recapture, data-availability, and nine CTP implementation years were checked independently for every scenario.
Detect a window with no usable GT estimate Implemented. The CTP stops rather than silently inventing an estimate. Long or clustered gaps need deliberate failure-handling rules for a full MSE.
Apply one reduced release and harvest sample size to all future GT years Implemented. Supports a constant lower-effort sensitivity.
Apply different effort levels in different future years Implemented and tested. Release and harvest sample-size vectors are keyed by release year. The observed 2025–2026 pipeline is shared by all four long-horizon arms before the 2027 programme scenarios diverge.
Simulate a future CPUE-q failure or step change Not exposed as a projection scenario. The current path carries forward fitted catchability and creep. The optional joint-stress experiment needs a defined q trajectory and code support.
Refit the full stock assessment at future assessment years Not implemented in run_projections(). It projects from a fixed accepted posterior and runs the CTP’s internal model, but does not update the full assessment posterior. Experiment A must use separate refits. A fully closed-loop assessment experiment would be a larger MSE extension.
Keep all non-GT stochastic draws paired when GT rows are removed Implemented and tested. CPUE, GT, HSP, and POP use component-specific streams, and retained GT rows use stable release-year keys. Use the same posterior rows, component seeds, and non-GT inputs in both arms.
Retain weak-stock draws when nominal TAC cannot be taken Implemented, tested, and independently audited. The opt-in operating-model rule scales every within-year allocation proportionally to keep raw harvest at or below 0.9, while retaining nominal advice and realized catch separately. The package default remains fail-closed. Report the probability and magnitude of catch shortfall; do not hide infeasible advice by dropping draws, and interpret stock risk alongside the lower realized removals.
Project far enough to see spawning-stock consequences Implemented in the long-horizon comparison. Dynamics extend through 2055 and the population state through 2056. This spans nine CTP updates and more than the documented 10–15+ year pathway from an age-2 cohort to spawning-stock outcomes.

The code now contains the calendar, year-specific effort, common-posterior, and random-stream machinery for the four-arm long-horizon comparison. This is still an assessment-posterior-conditioned closed-loop CTP analysis, not a formal revalidation or retuning of the MP across the full CCSBT operating-model reference set.

Proposed stages

Stage 0: agree the question

  • Decide whether the immediate question is CTP robustness, value to the stock assessment, a cost comparison, or preparation for a revised MP.
  • Confirm the latest operational GT schedule and map release years to data availability.
  • Agree whether results are internal screening material or formal ESC output.

Stage 1: make the projection comparison identifiable

Stage 2: small diagnostic screen

  • Use this screen to decide whether to retain the scenario and refine the horizon and performance measures.

Stage 3: assessment sensitivity

Stage 4: production analysis or full MSE

  • Add the CPUE-q joint stress only if it has a defensible specification.
  • If the question becomes a change to the CTP or long-term program design, move from screening projections to a formal MSE and MP-review process.

Questions outside this comparison

  • The analysis does not attach monetary costs to the four scenarios.
  • A formal MP review must decide which operating-model reference set, low- recruitment stresses, performance measures, and tuning criteria are needed for full MSE revalidation.
  • A CPUE-q joint stress remains separate unless evidence defines its magnitude, start year, and duration.
  • Any proposed programme change still requires scientific review through the exceptional-circumstances and meta-rule process.

Working provenance

  • OMMP16 report, projection workplan paragraph 54 (Commission for the Conservation of Southern Bluefin Tuna 2026).
  • Preece, Davies, Galeano, Hillary, and Eveson (2025), especially sections 3–6 (Preece et al. 2025).
  • Internal email thread, Exploring Gene Tagging Alternatives, 29 April–13 May 2026.
  • Internal email thread, 2026 MP evaluation - need a breakdown to justify potential budget, August 2025.
  • Darcy Webber–Ana Parma discussion, 23–24 July 2026.

The correspondence entries document planning context only. Formal scientific claims and any future terms of reference should be tied to approved CCSBT documents.

Earlier historical-omission comparison

The earlier run retains the accepted ESC31 base fit as the reference and changes GT availability in both the alternative historical assessment and its future monitoring schedule. The base MLE was not refitted. The first gate was an alternative MLE using the same model configuration, priors, parameter map, optimizer, and biological and numerical acceptance criteria. That MLE, its posterior, and the paired diagnostic projection are now complete.

The posterior review below uses all 3,000 retained draws from each fit. The initial projection comparison uses 100 chain-balanced posterior draws from each fit. It is retained as a diagnostic of historical information loss, not as the programme recommendation or the main future-monitoring comparison. Its different historical posteriors are specifically avoided in the common- posterior long-horizon analysis below.

Selected alternate-cohort design

The historical-omission experiment has one implementation: odd_release_cohorts. It drops the complete GT estimate for releases in 2017, 2019, 2021, and 2023, including the associated recapture work one year later. It retains the four GT estimates from releases in 2016, 2018, 2022, and 2024.

This is the requested every-second-estimate comparison. Its operational calendar is important: retained even-year releases are recaptured in the following odd calendar year. It therefore does not represent a complete shutdown of all fieldwork in odd calendar years. A literal odd-calendar-year shutdown would remove the recaptures for every retained even-year release and, with the current one-year release–recapture lag, leave no usable GT estimates.

Reproducible run contract

  • Fit artifacts are written under ESC31/runs/gtskip/; failed MCMC attempts and their diagnostics are retained rather than overwritten invisibly.
  • The first review gate stops after the alternative MLE; no MCMC or projection is launched until that fit is accepted for further work.
  • The alternative MLE must have convergence code zero, maximum gradient no greater than 0.01, estimability, valid biological states, exact catch accounting, and the accepted LL4 configuration.
  • The alternative MCMC uses four chains, 150 warmup and 750 retained iterations per chain, dense metric, the MLE mode as the chain start, adapt_delta = 0.999, maximum treedepth 13, and seed 73015.
  • Posterior acceptance requires maximum rank-normalized R-hat below 1.01, bulk and tail ESS at least 400, no divergences, no maximum-treedepth hits, and valid biological states for every retained draw.
  • The base and alternative 100-draw projections use the same retained iteration numbers in each chain and common component-specific random streams. CPUE, GT, HSP, and POP streams are isolated so removing GT years cannot shift the HSP or POP random draws.
  • Projected skip controls distinguish release years from harvest-sampling or recapture years. The generated calendar and the GT rows reaching each CTP update will be saved and reported.
  • The fit, MCMC, and paired projection artifacts are retained separately from the accepted production assessment and projection caches.

Only files carrying the odd_release_cohorts experiment name are inputs or outputs of the historical-omission analysis documented here. The older unprefixed gtskip_mle_* files and esc31_gtskip_odd_calendar_fieldwork.sbt.rds are preserved exploratory artifacts for different designs; they are not accepted results and must not be substituted for the named files. The local artifact registry in runs/gtskip/README.md records that distinction.

MLE review

ImportantFirst MLE gate complete

The alternate-release-cohort alternative has been fitted to maximum likelihood and passes every numerical, estimability, biological-state, catch-accounting, harvest-wall, and LL4 gate. Its posterior and paired diagnostic projection are reviewed separately below.

The executable comparison retains four of the eight historical GT observations. An independent input audit confirms that the numerical contents of all other model data, the parameter map, priors, bounds, and optimizer controls match the accepted base. Table 4 reports the acceptance checks. The alternative was initialized from the base MLE, but every active parameter was then re-estimated.

Table 4: Numerical and biological acceptance checks for the base and alternate-cohort maximum-likelihood fits.
Check Requirement Base Alternate-cohort GT MLE
Historical GT observations 8 versus 4 8 4
Final negative log-likelihood Finite 7616.637822 7600.761697
Optimizer convergence code 0 0 0
Maximum absolute gradient ≤ 0.01 2.60 × 10⁻¹⁰ 5.72 × 10⁻⁵
Estimability Pass Pass Pass
Biological-state gate Pass Pass Pass
Maximum raw seasonal harvest at age ≤ 0.9 0.617533 0.609524
Minimum populated numbers-at-age > 0 1696.516 1723.101
Maximum scaled catch error ≤ 10⁻⁸ 2.53 × 10⁻¹⁶ 3.65 × 10⁻¹⁶
Harvest-wall objective contribution Report (not a gate) 2.57 × 10⁻²¹ 5.17 × 10⁻²²
Continuation contribution ≤ 10⁻⁸ 0 0
LL4 standard-fleet contract Pass Pass Pass

The alternative objective is 15.876 units lower, but this is not evidence of a better-fitting model: its objective omits the four GT likelihood contributions from the omitted release cohorts. The useful comparison is the change in fitted stock quantities in Table 5, not the raw objective difference.

Table 5: Maximum-likelihood stock and parameter quantities for the base and alternate-cohort gene-tagging fits.
MLE quantity Base Alternate-cohort GT Difference Relative difference
B₀ 6,422,189 6,426,113 +3,924 +0.06%
2025 TRO 1,653,497 1,777,381 +123,884 +7.49%
2025 relative TRO 0.2575 0.2766 +0.0191 +7.43%
2025 recruitment 3,216,124 3,405,529 +189,405 +5.89%
2026 TRO 1,717,909 1,860,853 +142,944 +8.32%
2026 relative TRO 0.2675 0.2896 +0.0221 +8.25%
2026 recruitment 3,645,944 3,878,668 +232,723 +6.38%
M₀ 0.36574 0.36969 +0.00395 +1.08%
M₄ 0.16167 0.16342 +0.00175 +1.08%
M₁₀ 0.10736 0.10852 +0.00116 +1.08%
M₃₀ 0.45771 0.45994 +0.00223 +0.49%
CPUE q 0.95749 0.96826 +0.01078 +1.13%
Figure 1: Base and alternate-cohort maximum-likelihood trajectories from 2000 through the start-of-2026 model state. Points show every annual estimate. These are fitted MLE trajectories and contain no posterior uncertainty.

Retaining every second GT estimate still raises the recent fitted recruitment peaks and the terminal stock trajectory in this MLE (Figure 1). The largest proportional recruitment change over the fitted trajectory is about 38.9% in 2015; by 2026, recruitment is 6.4% higher and relative TRO is 0.0221 higher. The posterior review below evaluates uncertainty around the assessment trajectories, and the 100-draw screen then evaluates the first CTP response. The MLE result alone is not a management conclusion. The accepted ESC31 base, grid, and production projections remain closed while this separate robustness analysis continues.

The accepted fit is runs/gtskip/esc31_gtskip_odd_release_cohorts.sbt.rds. Reproducible review tables and the plotted trajectories are stored beside it as gtskip_odd_release_cohorts_mle_diagnostics.csv, gtskip_odd_release_cohorts_mle_management.csv, and gtskip_odd_release_cohorts_mle_review.rds.

Alternate-cohort posterior review

ImportantPosterior gate complete

The local four-chain alternate-cohort MCMC passed every numerical and biological-state acceptance gate. It is an accepted assessment sensitivity posterior for this screen, not a GT-skipping projection or management procedure result.

The sampler used the pre-specified dense-metric contract: four chains, 150 warmup iterations and 750 retained iterations per chain, exact MLE-mode starts, adapt_delta = 0.999, maximum treedepth 13, and seed 73015. Table 6 records the resulting diagnostics. The initial automatic step-size search had a few warmup-only divergences; the 3,000 retained draws used for inference had none.

Table 6: Sampling and biological-state diagnostics for the accepted alternate-cohort gene-tagging posterior.
Check Requirement Alternate-cohort posterior
Chains 4 4
Warmup per chain 150 150
Retained per chain 750 750
Total retained draws 3,000 3,000
Maximum rank-normalized R-hat < 1.01 1.0076
Minimum bulk ESS ≥ 400 1,632.1
Minimum tail ESS ≥ 400 1,153.3
Divergences after warmup 0 0
Maximum-treedepth hits 0 0
State draws expected / checked 3,000 / 3,000 3,000 / 3,000
Invalid or non-finite state draws 0 0
Maximum raw seasonal harvest at age ≤ 0.9 0.8731
Minimum populated numbers-at-age > 0 603.348
Maximum continuation penalty ≤ 10⁻⁸ 0
Overall posterior gate Pass Pass

Historical GT rows

The alternate fit skipped releases in 2017, 2019, 2021, and 2023 and their associated recapture years 2018, 2020, 2022, and 2024. It retained releases in 2016, 2018, 2022, and 2024, recaptured respectively in 2017, 2019, 2023, and 2025. This is the linked cohort-level schedule in Table 7; it is not a blanket shutdown of all work in odd calendar years.

There is no 2020 release row in Table 7 because tagging was cancelled in that field year. That source-data gap is separate from the retained/omitted pattern imposed by this sensitivity.

Table 7: Historical gene-tagging data and the rows retained or omitted in the alternate-cohort assessment fit.
Release year Release age Recapture year Releases Scanned samples Matches Alternate-fit status
2016 2 2017 2,952 15,389 20 Retained
2017 2 2018 6,480 11,932 67 Omitted
2018 2 2019 6,295 11,980 66 Retained
2019 2 2020 4,242 11,109 31 Omitted
2021 2 2022 6,401 10,742 41 Omitted
2022 2 2023 5,084 14,714 38 Retained
2023 2 2024 2,759 13,297 14 Omitted
2024 2 2025 3,522 11,011 11 Retained

Posterior trajectories and GT fits

Figure 2 compares the assessment trajectories using all 3,000 retained draws from each posterior. These are separately sampled base and alternate assessment distributions; they are not paired projection draws.

Figure 2: Base and alternate-cohort posterior trajectories from 2000 through the start-of-2026 model state. Lines are posterior medians and ribbons are equal-tailed 95% credible intervals from all 3,000 retained draws in each fit.

At 2026, median relative TRO is 0.3044 in the alternate fit versus 0.2817 in the base, an absolute difference of 0.0227 (8.0% relative to the base median). The corresponding 95% intervals are 0.2447–0.3829 and 0.2216–0.3526. Median recruitment is 3.819 million versus 3.680 million, a difference of 0.139 million (3.8%); its two wide 95% intervals also overlap substantially. The upward terminal-TRO direction seen at the MLE therefore remains in the posterior medians, but the assessment uncertainty is material and this is not a paired test of projected outcomes.

The fitted GT expectations in Figure 3 use the sample sizes in Table 7. Orange points are observations used in the corresponding likelihood. Grey crosses locate the four observations omitted from the alternate fit; they are shown only for context and did not contribute to its likelihood.

Figure 3: Observed and posterior fitted historical GT match counts for the base and alternate-cohort assessments. Blue open points are posterior medians and vertical blue intervals are equal-tailed 95% credible intervals for expected matches. Grey crosses in the alternate panel are omitted observations, not fitted data.

Paired 100-draw CTP projection screen

ImportantDiagnostic projection gate passed

The annual-GT reference and alternate-release-cohort arms each completed 100 chain-balanced draws. All 200 historical states and all 400 CTP update optimisations across the two arms passed their configured gates. This accepts the run as a diagnostic screen; it does not promote it to the 2,000-draw production assessment or make it ESC advice.

The comparison changes two linked parts of the workflow: the alternate arm starts from the accepted historical-omission posterior reported above and also omits the corresponding future GT release cohorts. It therefore represents a continued alternate-cohort program, not a projection-only test in which both arms start from the same posterior.

The 100 rows comprise 25 retained iterations from each of four chains. Both arms use draw seed 44, projection seed 102, the same retained iteration numbers, the same recruitment and selectivity rules, the same fixed TAC and fleet allocation, and component-specific common random streams. The projection covers model and monitoring years 2022–2035 and the start-of-2036 population state.

Future GT calendar

The future schedule in Table 8 skips releases in 2025, 2027, 2029, 2031, and 2033 and their linked recaptures in 2026, 2028, 2030, 2032, and 2034. Retained releases still require recapture work in the following odd calendar year.

This retained 100-draw screen is deliberately counterfactual from release year 2025: it skips the 2025 release even though that field programme has already occurred. It is therefore not the current operational pipeline. It also omits odd release years, whereas the later common-posterior GT2 and GT3 arms retain 2027 and omit even release years from 2028. “Alternate cohort” describes the frequency in both experiments, not a shared parity or calendar.

Table 8: Future release and recapture calendar used in the alternate-release-cohort projection.
Release year Recapture year Alternate-arm GT estimate
2025 2026 Skipped
2026 2027 Retained
2027 2028 Skipped
2028 2029 Retained
2029 2030 Skipped
2030 2031 Retained
2031 2032 Skipped
2032 2033 Retained
2033 2034 Skipped
2034 2035 Retained

Projection diagnostics

Table 9 records the complete numerical gate. The alternate arm’s maximum CTP gradient, (4.9811^{-4}), passes the (5^{-4}) tolerance but is close enough that it should be watched in a larger run. Every final CTP Hessian was positive definite and every historical state passed.

Table 9: Numerical and biological-state diagnostics for the paired 100-draw CTP projection screen.
Check Requirement Annual GT Alternate release cohorts
Projection draws expected 100 100 100
Historical states checked 100 100 100
Historical states passed 100 100 100
CTP updates expected 200 200 200
CTP updates passed 200 200 200
Positive-definite CTP Hessians 200 200 200
Maximum absolute CTP gradient ≤ 0.0005 0.0004668 0.0004981
Elapsed time Report only 144.8 s 144.9 s

CTP inputs and TAC response

The first CTP update in 2030 can use GT information through release year 2026; the second, in 2033, can use information through 2029. Table 10 shows that skipping estimates did not simply lower the weighted GT abundance. The remaining match-weighted estimates produced a slightly higher median at both updates, and the median GT multiplier stayed at its neutral value of one in both arms.

Table 10: GT inputs and TAC outcomes at each CTP update in the 100-draw screen. Counts in the final columns are numbers of projection draws.
CTP implementation year Latest GT release year Median GT abundance: annual Median GT abundance: alternate Median GT multiplier: annual / alternate Median TAC: annual Median TAC: alternate Decrease / unchanged / increase: annual Decrease / unchanged / increase: alternate
2030 2026 2,244,647 2,288,038 1.00 / 1.00 26,647.0 26,647.0 0 / 0 / 100 0 / 0 / 100
2033 2029 1,953,125 2,222,222 1.00 / 1.00 26,802.3 26,818.8 1 / 37 / 62 2 / 30 / 68

Figure Figure 4 shows the resulting marginal TAC paths.

Figure 4: Nominal total TAC for the annual-GT and alternate-release-cohort arms. Lines show medians and ribbons show empirical 95% intervals across 100 draws. Dashed vertical lines mark CTP update years; TAC is fixed before 2030.

The central TAC paths are very similar, but the result is more nuanced than “no effect”:

  • The 2026–2029 TACs are prescribed and therefore identical.
  • In 2030, the median unconstrained TAC is about 35,542 t in the annual arm and 35,502 t in the alternate arm. Every draw reaches the 3,000 t increase limit, so both implemented medians are exactly 26,647 t through 2032.
  • At the 2033 update, the marginal medians differ by only 16.5 t (26,802.3 versus 26,818.8 t). Their empirical 95% intervals are 26,647–29,647 t and 26,393–29,647 t, respectively.
  • The draw-level responses are not identical: 74 of 100 paired computational rows differ at the 2033 update, and the empirical paired difference interval is -2,852 to +2,540 t. This artificial common-random-number pairing is useful for a screen, but it is not a credible interval from a formally joint posterior.
  • Median cumulative nominal TAC over 2026–2035 is 253,960 t for annual GT and 253,998 t for the alternate arm. The paired median difference is zero, while its empirical 95% interval is -8,557 to +7,619 t.

The first update is therefore uninformative about the operational effect of GT gaps because the TAC-change cap binds. The second update suggests very similar central TAC advice but enough draw-level variation that equivalence should not be claimed from 100 draws.

Projected stock status and risk

Figure Figure 5 and Table Table 11 report the corresponding stock-status screen.

Figure 5: Projected relative total reproductive output for the annual-GT and alternate-release-cohort arms. Lines show medians and ribbons show empirical 95% intervals across 100 draws. The dashed horizontal line is the relative-TRO risk threshold of 0.2.
Table 11: Selected relative-TRO summaries from the paired 100-draw projection screen.
State year Annual GT median (95% interval) Alternate median (95% interval) Difference in marginal medians
2030 0.330 (0.265–0.421) 0.363 (0.280–0.461) +0.034
2033 0.345 (0.262–0.498) 0.387 (0.289–0.516) +0.042
2036 0.344 (0.241–0.571) 0.384 (0.270–0.546) +0.040

The alternate arm retains the higher terminal stock-status direction already seen in its historical assessment posterior. No draw in either arm falls below relative TRO 0.2 during the configured short-term window (2030–2032) or longer window (2033–2036). With 100 draws this means only that the observed screening frequency was zero; it does not establish a zero risk probability.

The higher alternate trajectory must not be attributed solely to future GT skipping. It also reflects conditioning on the different accepted historical posterior, and the intervals overlap substantially. Moreover, the current horizon is too short to capture the full documented 10–15+ year path from a future age-2 cohort to spawning-stock outcomes.

Interpretation and next gate

The screen supports three provisional comments:

  1. Removing alternate GT estimates did not automatically depress the CTP’s weighted GT input; which cohorts remain and their match weights matter.
  2. The median TAC path was effectively unchanged through the first update and differed by only 16.5 t at the second, largely because the first update hit the TAC-change cap and the median GT multiplier remained neutral.
  3. Individual draw responses and cumulative differences were much wider than the marginal medians suggest, while the stock-status comparison also includes historical-posterior displacement.

Those limitations motivate the common-posterior long-horizon comparison below. The earlier result remains useful for showing that historical GT omission can move the assessment posterior, but its 16.5 t median TAC difference is not used to judge whether a future programme is acceptable.

Common-posterior long-horizon closed-loop comparison

This comparison isolates future monitoring. Every arm uses the same 2,000 chain-balanced draws from the accepted base posterior, the same future recruitment and selectivity paths, the same catch allocation and removal assumptions, and component-specific common random numbers. Only the future GT effort and availability schedule changes.

The projection covers dynamics from 2022 through 2055 and the resulting state through 2056. Fixed TACs apply from 2026 through 2029, after which the CTP is run at nine implementation years: 2030, 2033, 2036, 2039, 2042, 2045, 2048, 2051, and 2054. This removes the two-update horizon limitation of the earlier screen and spans the 10–15+ year pathway from juvenile recruitment to the spawning stock.

Nominal CTP TAC advice and realized catch are retained separately. No posterior draw is discarded or replaced if its stock becomes too small to take the nominal TAC. Instead, an explicit operating-model implementation rule reduces all seasonal and fishery allocations in that draw-year by one common factor until the raw combined seasonal harvest rate is no greater than 0.9. The report therefore audits catch shortfall as an outcome. Because this rule can reduce late-horizon removals in weak-stock draws, risk results must be interpreted together with both nominal TAC and realized catch rather than as an unqualified consequence of the monitoring design.

The programme’s target CV of 0.25 is retained as an operational design target, not imposed as a model acceptance gate. The results report simulated match counts and use 16 matches (the crude Poisson quantity (1/0.25^2)) only as a descriptive precision proxy. That proxy is not a replacement for a formal design-based CV calculation.

NoteScope of this closed-loop result

This is a closed-loop, posterior-conditioned CTP simulation: monitoring data are generated from each operating-model path, passed to the CTP, and the resulting TAC changes feed back into subsequent population and observation paths. It is substantially more informative than the 100-draw screen. It is not, however, formal revalidation or retuning of the MP across the complete CCSBT operating-model reference set, and it cannot authorize a change from the annual programme.

ImportantProduction comparison complete; retain annual gene tagging

All four 2,000-draw arms passed the independent numerical and biological review. Reducing the annual harvest sample lowered simulated match precision but retained five recent estimates in every CTP window and produced results very close to the full-effort reference in this experiment. Deliberate gaps reduced the recent-information window to two or three estimates, and the added failure reduced it to one estimate at the 2036 CTP calculation. Those gaps increased TAC variability and gave slightly poorer long-horizon stock-risk indicators, although the paired simulation intervals were broad and included zero. These counterfactuals do not provide a basis for reducing the proposed annual programme.

Independent validation

The production run reused the same 2,000 chain-balanced posterior rows and paired non-GT stochastic inputs in all arms. Each arm passed all 2,000 historical-state checks and all 18,000 CTP optimisations. Every final CTP Hessian was positive definite, the maximum absolute gradient met the pre-specified \(5 \times 10^{-4}\) tolerance, populated abundance remained positive, raw harvest reached but did not exceed the configured 0.9 ceiling beyond its \(10^{-10}\) numerical tolerance, and the feasibility penalty was effectively zero. Table Table 12 records these gates.

Table 12: Independent numerical and biological review of the four long-horizon programme arms.
Programme scenario Historical states passed CTP updates passed Maximum CTP gradient Minimum Hessian eigenvalue Overall gate
Annual full effort 2,000 / 2,000 18,000 / 18,000 0.000499976 12.025 Pass
Annual reduced effort 2,000 / 2,000 18,000 / 18,000 0.000499976 11.897 Pass
Deliberate alternate-cohort skips 2,000 / 2,000 18,000 / 18,000 0.000499930 12.051 Pass
Deliberate skips plus unplanned failure 2,000 / 2,000 18,000 / 18,000 0.000499953 12.034 Pass

Monitoring coverage and simulated precision

Figure Figure 6 and Table Table 13 separate the number of usable cohort estimates from the precision of those retained.

Figure 6: Median number of usable GT estimates in the CTP’s five-cohort window. Annual full and annual reduced effort retain all five estimates. Deliberate skips alternate between two and three; adding the unplanned 2029 failure leaves only one usable estimate at the 2036 calculation.

Reducing annual harvest sampling from 13,000 to 10,000 lowered the median simulated matches per active cohort from 29 to 22 and more than doubled the mean frequency below the descriptive 16-match precision proxy. It did not remove an estimate from the CTP window. The skip scenarios retained the full-effort precision for cohorts that were actually sampled, but had far fewer cohort estimates available. Match precision for an observed cohort and coverage of recent cohorts are therefore distinct consequences.

Table 13: Simulated precision of available future GT estimates and coverage of the CTP’s recent five-cohort window. The 16-match quantity is a descriptive Poisson proxy for CV 0.25, not a formal design-based precision calculation.
Programme scenario Median matches per active cohort Probability below 16-match proxy Minimum median usable estimates per CTP window
Annual full effort 29 10.9% 5
Annual reduced effort 22 23.2% 5
Deliberate alternate-cohort skips 29 10.9% 2
Deliberate skips plus unplanned failure 29 10.9% 1

TAC response

Figures Figure 7 and Figure 8 show marginal and paired TAC responses; Table Table 14 summarizes the paired cumulative differences.

Figure 7: Nominal total TAC under the four programme scenarios. Lines are marginal medians and ribbons are empirical 95% simulation intervals across 2,000 draws. Dotted vertical lines mark the nine CTP implementation years.

All arms are identical through the capped 2030 response. Later marginal median TACs are higher under deliberate gaps, especially after the added failure, but this is not evidence that missing monitoring improves performance. Missing cohorts change the recent match-weighted GT input in either direction rather than imposing a downward adjustment. The median paired cumulative increase is 1,543 t with deliberate skips and 3,332 t after the added failure, only about 0.2% and 0.4%, respectively, of the annual-full median cumulative TAC of 819,398 t. Their paired 95% intervals are much wider and span zero. The higher central nominal advice is also accompanied by slightly lower central stock status and higher long-term risk.

Figure 8: Paired nominal-TAC differences from annual full effort. Lines are mean paired differences and ribbons are empirical paired 95% simulation intervals. The paired median is zero at many update years: initially most draws share the same capped response, while later non-zero differences increasingly straddle zero even though most paired draws differ.
Table 14: Paired cumulative nominal-TAC differences from the annual full-effort reference over 2026–2055. Intervals describe the simulated paired outcomes and are not posterior credible intervals for a formally revalidated MP.
Programme scenario Median paired cumulative TAC difference (t) Empirical paired 95% interval (t) Mean absolute paired difference (t) Draws with a different cumulative TAC
Annual reduced effort 0 -16,116 to 16,911 4,059 73.8%
Deliberate alternate-cohort skips +1,543 -43,030 to 52,483 14,348 86.0%
Deliberate skips plus unplanned failure +3,332 -58,170 to 72,301 20,413 88.2%

Stock status and risk

Figure Figure 9 and Table Table 15 show the long-horizon stock distributions and the associated risk summaries.

Figure 9: Projected relative total reproductive output under the four programme scenarios. Lines are marginal medians and ribbons are empirical 95% simulation intervals. The dashed horizontal line is the relative-TRO risk threshold of 0.2.

The annual reduced-effort arm remains very close to full annual effort. The two gap scenarios have progressively lower 2056 median relative TRO and higher full-horizon probabilities of ever falling below 0.2. The central changes are small relative to the simulation uncertainty: even the combined skip-plus- failure arm has a paired 2056 difference interval of -0.0672 to +0.0588. Separation appears mainly in the final years, consistent with the delay from age-2 recruitment information to spawning-stock consequences.

Table 15: Long-horizon stock status and risk under the four programme scenarios. The paired intervals reflect common-posterior simulation differences; they are not evidence that the scenarios are equivalent.
Programme scenario 2056 relative TRO, median (95% interval) Paired median difference from annual full (95% interval) P(2056 relative TRO < 0.2) P(ever < 0.2), 2030–2056
Annual full effort 0.291 (0.046–0.637) Reference 23.9% 26.6%
Annual reduced effort 0.290 (0.047–0.640) 0.0000 (-0.0154 to +0.0146) 23.8% 26.7%
Deliberate alternate-cohort skips 0.287 (0.047–0.636) -0.0011 (-0.0478 to +0.0400) 25.3% 28.4%
Deliberate skips plus unplanned failure 0.283 (0.043–0.638) -0.0026 (-0.0672 to +0.0588) 26.3% 29.4%

In the first decade (2030–2039), the probability of ever falling below 0.2 is about 2.8% in every arm. In the final 2050–2056 window it rises from 25.6% under annual full effort to 27.4% with deliberate skips and 28.6% with the additional failure. Across the complete 2030–2056 horizon, the corresponding ever-below risk rises from 26.6% to 28.4% (+1.9 percentage points) and 29.4% (+2.9 percentage points), using the underlying unrounded probabilities (1.85 and 2.85 percentage points before rounding). The probability of being below 0.2 specifically in 2056 rises by 1.4 and 2.4 percentage points. This analysis therefore detects a modest adverse long-term direction from missing monitoring, not a sharply separated short-term effect; the paired final-stock intervals still span zero.

Nominal advice and realized catch

Figure Figure 10 shows where nominal TAC advice cannot be fully realized under the operating-model harvest ceiling.

Figure 10: Probability that nominal TAC cannot be fully realized under the operating- model harvest ceiling. The event remains rare but becomes more common late in the projection as uncertainty in weak-stock paths expands.

Median cumulative catch shortfall is zero in every arm, and only 2.5–2.6% of draws experience any shortfall. The 97.5th percentile of cumulative shortfall is about 7 t under either annual scenario, 389 t under deliberate skips, and 1,215 t after the added failure. A few weak-stock paths have much larger annual shortfalls. Stock-risk results must therefore be interpreted alongside realized catch: the retained weak-stock draws are not allowed to remove an infeasible nominal TAC.

Scientific interpretation

  1. Lower annual sample size primarily affects precision. It retains five recent GT estimates in each CTP window but roughly doubles the frequency of active cohorts falling below the descriptive match-count proxy. In this isolated GT-only comparison, its TAC and stock-risk results remain close to annual full effort.
  2. Skipping cohorts primarily affects coverage. Regular gaps reduce the five-cohort window to two or three estimates. An adjacent unplanned failure briefly leaves only one, increasing TAC dispersion and producing the largest adverse central stock-risk changes.
  3. The direction is more informative than the apparent separation. The gap scenarios have higher central nominal TACs and slightly lower relative TRO, but paired intervals span zero and the long-horizon stock distributions overlap strongly. Neither equivalence nor a precise causal effect should be claimed from this one posterior-conditioned experiment.
  4. Annual GT remains the appropriate reference. The comparison excludes a CPUE-q failure, future full-assessment refits, alternative operating-model reference sets, and MP retuning. Any proposed reduction still requires the formal MSE, MP-review, exceptional-circumstances, and meta-rule process.

The independent summaries and figures were generated with scripts/summarise-esc31-gtprogramme-mse.R from the atomic four-arm comparison and its four scenario caches. The review binds the exact input artifacts by MD5 and stops if calendars, paired non-GT inputs, biological states, CTP optimisations, or catch accounting fail their configured gates.

References

Commission for the Conservation of Southern Bluefin Tuna. 2026. Report of the Sixteenth Operating Model and Management Procedure Technical Meeting. OMMP16 meeting report. Commission for the Conservation of Southern Bluefin Tuna.
Preece, A., C. Davies, D. Galeano, R. Hillary, and P. Eveson. 2025. Impacts of Changes to the Gene-Tagging Program. CCSBT-ESC/2508/20. Commission for the Conservation of Southern Bluefin Tuna. https://www.ccsbt.org/system/files/2025-08/ESC30_20_AU_Implications_of_changes_to_GT_program.pdf.
Preece, A., and P. Eveson. 2026. Gene-Tagging Proposal 2027–2029. CCSBT-ESC/2608/13. Commission for the Conservation of Southern Bluefin Tuna.
Preece, A., P. Eveson, J. Hartog, et al. 2026. Update on the SBT Gene-Tagging Recruitment Monitoring Program 2026. CCSBT-ESC/2608/10. Commission for the Conservation of Southern Bluefin Tuna.