Review

Sign-off

The ESC31 base assessment, sensitivities, MCMC grid, MLE comparison grid, and production projections are accepted. The implementation defects found during review have been addressed in the assessment workflow and the sibling sbt package. There is no unresolved numerical or implementation defect that requires the base model, MCMC, grids, or projections to be rerun for ESC31.

The final base MLE and posterior pass their numerical and biological-state checks. All fourteen sensitivity MLEs are accepted, and all fourteen sensitivity posteriors are accepted under their recorded decisions. All nine MCMC grid cells, the balanced 2,000-draw posterior, the 108-cell direct-M MLE grid, and its 2,000-draw resample are complete. The four 2,000-draw production projection arms pass their reporting gates.

The latest sbt main-branch R CMD check, pkgdown, and benchmark workflows pass. The sbt2026 publication workflow runs the ESC31 static source checker and publishes the site successfully on every push to main.

Selected assessment specification

CPUE and composition weighting

The base reads the GAM22 B = 1000 raw annually varying CPUE CVs from CV_GAM22_20260713.xlsx and adds them to a fixed log-scale sigma of 0.20 in log-variance space. The resulting annual ordinary CVs range from 20.24% to 30.10%, with a mean of 21.34%.

The CPUE and composition likelihoods are intentionally downweighted. Most age- and longline-length-composition OSA SDNRs are below one and do not, by themselves, imply that those data should receive greater weight. The larger CPUE length-composition residuals on the separate package-default comparison are documented on the Data weighting page and are not ESC31 base statistics.

Natural mortality and harvest feasibility

The selected base uses M_switch = 2 and get_M_length(). M10 and M30 are estimated on the log scale, the length exponent is fixed at mc = -1, and M0 and M4 are derived from length at age. Every accepted posterior draw has finite, strictly positive mortality at every age.

All fisheries use standard selectivity-based catch conditioning. Observed total catch is reproduced exactly, while selectivity allocates removals by age. Accepted states require positive abundance, raw combined seasonal harvest at age no greater than 0.9, exact catch accounting, and zero continuation penalty. The preventive harvest wall has strength 10, onset 0.85, ceiling 0.90, and scale 0.01. It changes the objective but not catch or population dynamics.

LL3 and LL4

LL3 and LL4 remain separate fisheries. LL4 has one time-invariant selectivity block beginning in 1953, with estimated effects over ages 8–21 and terminal- age extension beyond age 21. Its fixed year correlation, age correlation, and selectivity sigma match LL3 at 0.5, 0.5, and 0.75. Length compositions remain soft likelihood information, and total catch remains exact.

The alternative direct-removal treatment is rejected. Its archived fit approached the harvest boundary, with raw harvest about 0.90, and retained a poor maximum gradient of about 1.47. It is unsuitable for production MCMC and is retained only as rejection and regression evidence until the unused code is removed in a later package-maintenance change.

Accepted numerical evidence

Component Accepted result
Base MLE Objective 7616.63782197, maximum gradient 2.59898e-10, estimable Hessian, maximum raw harvest 0.6175334, minimum populated abundance 1696.516, and maximum relative catch error 2.53e-16
Base MCMC Maximum R-hat 1.008701, minimum bulk ESS 1157.424, minimum tail ESS 1178.248, no divergences, no maximum-treedepth hits, and all 3,000 retained states biologically valid
Posterior MSY All 222,000 draw-year cells finite and biologically valid, with both optimisation objectives negative
Sensitivities Thirteen accepted MLEs and posteriors; NoPOPHSP alone uses the explicit two-divergence acceptance decision described below
MCMC grid Nine accepted cells and a balanced 2,000-draw posterior; maximum R-hat across cells is below 1.01, minimum bulk and tail ESS exceed 990, and there are no sampler pathologies
MLE grid The direct-M design contains 108 distinct fits and supplies a validated 2,000-draw full-objective-weighted resample
Production projections Four accepted 2,000-draw arms covering both allocation scenarios, NoUAM, and the direct-M MLE-grid resample
Independent CTP audit All 16,000 production updates reproduced by a separately written implementation within the stated tolerances

The base run signatures are f446f9eebca6b87e4b2e5834f81b616a for the MLE and 585dc82abe8232ac4f2bb8e6653d413c for the accepted posterior.

OMMP16 cross-check

Requirement ESC31 disposition Status
POP age uncertainty and revised conventional-tag mixing Implemented in the current likelihoods Done
Indonesian selectivity blocks The terminal block is 2021–2025; the 1976 node is deliberately retained Done
Age-4+ CPUE fleet and weighted length compositions Implemented as Fishery 7 with separate abundance and length-frequency selectivity selectors Done
Japanese NCNM/UAM series Included in base historical catches; the NoUAM sensitivity removes it before rebuilding data Done
Selectivity hyperparameters The confirmed ESC31 fixed values are retained; alternative generic prior estimation was not selected for the base Accepted base decision
CPUE CV treatment The selected base combines fixed log-SD 0.20 with raw annually varying GAM22 CVs; the constant-20% sensitivity is also reported Done
Composition likelihoods and sample sizes Multinomial likelihoods and the retained fleet-specific sample sizes are used Done
Posterior and deterministic grids Nine-cell MCMC grid, balanced 2,000-draw posterior, and 108-cell direct-M MLE comparison are complete Done
Projection calendar, selectivity, recruitment, allocation, and TAC lags Implemented, tested, and represented in all four production arms Done
Named sensitivities The requested assessment sensitivities are represented, with the qualifications below Done
Gene-tagging programme comparison The assessment sensitivity and four-arm 2,000-draw long-horizon comparison are complete and support retaining annual gene tagging Done
Japanese/Korean index and raised catch-at-length comparisons Provider-approved source products and metadata are not available in the current assessment inputs Future external-data work

Recorded qualifications

NoPOPHSP

The canonical NoPOPHSP posterior contains two divergent transitions but passes the accepted R-hat, ESS, treedepth, MLE, and complete biological-state checks. Decision esc31_2026_no_pop_hsp_accept_two_divergences_12000_draws_v1 accepts those two divergences for this sensitivity only. A longer-warmup candidate removed the divergences, but maximum R-hat was 1.033873, minimum bulk ESS was 82.44045, and minimum tail ESS was 186.027, all for par_log_m10. The candidate was therefore rejected and did not replace the canonical result. No further sampler run is planned for ESC31.

Troll and conventional-tag diagnostics

The accepted Troll sensitivity adds the trolling index and retains the calibrated base aerial tau of 0.59. OMMP16 did not specify a larger numerical tau; any additional aerial downweighting would be a separately defined future sensitivity.

Conventional-tag residuals use the likelihood-matched compResidual::resDirM() fallback. RTMB oneStepPredict() cannot execute the required multivariate discrete case. This is a diagnostic-method decision and does not change the fit.

Japanese size-frequency duplicates

Twenty-three duplicated year-month-area-length keys occur in the 1996 Japanese size data. The corrected diagnostic sums those records before normalisation. The effect is confined to 1996 and is small: the largest absolute bin change is 0.000529, and total-variation distance is about 0.001346.

The checksum-bound accepted 1996 composition remains frozen for ESC31, so no current input, fit, grid, or projection changes. Duplicate-safe construction is required at the next planned data refresh and is tracked in sbt2026 issue 6.

Projection and CTP sign-off

The last completed production projections use cache format 8. A later format-9 code change stages not-yet-decided future catch as zero and records realised catch separately. In the accepted format-8 runs there was no catch shortfall, no harvest penalty, and maximum raw harvest was below 0.57. Future placeholder catch therefore could not alter observations available to an earlier CTP update, and the format-8 caches remain accepted without a production rerun. Future projection runs will use format 9.

The independent CTP audit does not load or call sbt. It uses separately written C++/TMB CKMR equations and separate calendar, data-adapter, optimiser, harvest-control-rule, TAC-bound, allocation, and removal code. All 16,000 paired updates pass. There are no TAC-bound branch mismatches or allocation- input differences, and every independently calculated Hessian is positive definite.

The production report projects population dynamics through 2035, reports the resulting state through 2036, uses a relative-TRO threshold of 0.2, and reports 2030–2032 separately from 2033–2036.

Remaining work

The following items do not block ESC31 sign-off and do not require another base fit, posterior, grid, or projection run.

Item Disposition
Japanese/Korean external data Obtain provider-approved Korean index inputs and fully raised Korean and New Zealand catch-at-length products, including units, coverage, uncertainty, raising factors, missing-data treatment, and provenance. Treat these as future comparison or sensitivity inputs rather than altering the accepted base.
Duplicate-safe Japanese size aggregation Apply sum-before-normalisation construction at the next data refresh, as tracked in issue 6. Do not refresh or replace the accepted ESC31 inputs solely for this small 1996 diagnostic difference.
Retire direct-removal code After ESC31, remove the unused implementation, compatibility branches, tests, and documentation in one package-maintenance change while preserving the archived rejection evidence.
Stochastic projected selectivity A numerical perturbation was not specified for ESC31. The accepted projections retain fitted selectivity through 2025 and use the deterministic 2016–2025 mean thereafter. Any stochastic method remains a future package decision.

Published evidence

  • Introduction describes the assessment inputs and workflow.
  • Base model reports the accepted MLE, base posterior, residuals, and posterior MSY.
  • Sensitivities reports all fourteen accepted sensitivity MLEs and posteriors, including the NoPOPHSP decision and the conventional-tag exclusion.
  • Grids reports the accepted MCMC grid, balanced posterior, and direct-M MLE grid and resample.
  • Projections reports the four accepted 2,000-draw production arms.
  • GT skipping reports the accepted assessment sensitivity and long-horizon monitoring-programme comparison.
  • B10+ and TRO age structure explains the different recent behaviour of relative B10+ and relative TRO.

The assessment source, frozen inputs, fit identities, cache manifests, and independent audit outputs retain the detailed provenance behind these pages. The official Cape Town Procedure specification was used for the independent implementation cross-check.