Borrow the pieces you haven’t ported yet.
The first JAX entry point already runs a case. It calls the remaining Fortran routines through F2PY—the bridge that lets Python call Fortran. The port works from the start, even while it is “cheating.”
The short version: first we built a bridge to the old Fortran program. Then we built a measuring stick. Then we finished one difficult path through JAX. After that, the work changed from figuring out how to port CLUBB into applying a recipe that already worked.
CLUBB is an atmospheric model. Its trusted implementation is written in Fortran.
JAX is a Python numerical system that can compile calculations for CPUs and accelerators.
Porting here means recreating a Fortran calculation in JAX while keeping its numerical behavior checkable.
A human touch is one human-written prompt during an agent task.
The recipe
Pick a stage to see what it was for, what it produced, and what the human/agent rhythm looked like. These are stages of the method, not six strictly separate time periods.
Watch the method
Start at the top, then follow one branch all the way down. Check the whole run after each piece. When you move sideways, shared helpers are already JAX—ready to use again.
Illustrated call tree; each box stands for a piece of code, not a literal file or the complete CLUBB call graph. Swipe across the diagram to see both trees.
The first JAX entry point already runs a case. It calls the remaining Fortran routines through F2PY—the bridge that lets Python call Fortran. The port works from the start, even while it is “cheating.”
Run the hybrid program and the original Fortran program with the same inputs after each replacement. The selected CLUBB cases check real callers, shared state, and outputs together. A new unit test for every routine is not a prerequisite: the full-run comparison is already there. Small unit tests can still stress-test edge cases the selected runs do not exercise.
Depth-first means finishing one call path all the way down before moving sideways. Different Fortran branches often use the same small routines. Once those helpers are ported, compiled, and checked, later branches can call their existing JAX versions. We reuse working code, not just lessons learned—and keep each change small enough to review.
The code and the checks around it
We count lines added in Git commits connected to the port. One view includes tests; the other leaves them out. Both include support scripts, configuration, comments, and blank lines.
Added lines are not final program size. Rewriting a line can count it again in a later commit. Formatting changes count too. Code that was never committed is outside this measure.
C1 includes tests; C2 excludes them. The chart selector applies to every code chart on this page. Documentation is counted separately, not included in either measure.
How people talked to the agents
Some short messages were simply “go” prompts—such as “continue,” “go ahead,” or “OK, go for it, do a good job”—rather than new instructions. A short prompt can still contain an important correction or decision.
What the human touches did
Length is one view; purpose is another. Each row is a stage, and each colored width is the share of its prompts assigned to that role.
Each prompt gets one category using wording-based rules. A message can do several things, so these labels are a guide to the conversation—not a precise judgment of intent or necessity.
The part we care about most
The first complete dependency path taught the rules. The next wave reused them across much more code.
The same-input full-run tests were already established. Each new branch could use that comparison, with additional cases and checks added as problems surfaced.
Markdown instructions and reusable agent commands captured how to port, run comparisons, and finish a change. Later tasks could refer to that workflow instead of restating every step.
JAX-compatible objects, array shapes, branching, and updates had been worked through on one path. Those patterns carried across branches, and already-ported leaf routines could be called again directly.
Real calendar time
Each bubble is one stage, placed at the middle of its prompt activity. Bubble size shows prompt count.
This shows how the size of committed changes per human touch evolved as the code and porting workflow matured. It is not a model leaderboard: stages had different scopes, and the chart does not control for model differences or reasoning settings. Formatting and repeated rewrites also count as added lines. The first core path and later core scale-out are the closest comparison here.
Inside each stage
The stages tell us what the work was for. These bars break down the code it changed: calculations, the driver around them, tests, and supporting tools.
Categories follow the files changed, not the reason for every individual line. One stage can touch many parts of the program. Documentation stays outside these code counts.
The work behind the text
Tokens are pieces of text processed or generated by a model—not lines of code. Input includes instructions, conversation history, code, and tool results, often read again on later calls. Output includes explanations, tool commands, code, and reasoning.
T1 = T2 + T4. Cached input is already in T2; reasoning is already in T4. Do not add them again. Large input counts include repeated context; they are not billions of newly written words.
A tiny example: 1,000 input tokens (including 800 cached) + 200 output tokens (including 50 reasoning) = 1,200 total. The subsets show how the tokens were used; they do not increase that total.
Token counts are not a subscription invoice or a count of code lines. Official token definitions ↗
Building on earlier work
Alex Connolly’s CLUBB-JAX gave us reusable numerical routines and a reference for other translations. Its clearest contribution came as we expanded the core beyond the first completed branch.
Comparisons with advance_xm_wpxp and advance_xp2_xpyp informed updates across whole atmospheric columns. We also copied parameter_indices.py, adapting its name-to-array-position lookup to Python’s zero-based indexing.
We copied 13 PDF-related files. PDF means probability distribution: variations in moisture, temperature, and motion within a grid cell. Some helpers carried over almost directly; larger routines needed expansion and integration.
56 source files at the June 16 core checkpoint: 49 core files plus seven supporting data-object files. Package markers, tests, the driver, and later microphysics are excluded.
Each file counts equally. Categories describe the June code and involve review judgment; they do not measure lines borrowed or hours saved. Percentages are rounded.
Each note distinguishes direct evidence, a review judgment, or an unresolved case.
What carried forward
Root solvers, matrix helpers, and many cloud-distribution formulas carried forward; the root and matrix calculation bodies still match in September. Around them, we expanded missing branches, connected statistics and data objects, and validated complete runs against Fortran.
calc_roots.py (polynomial equations) and matrix_operations.py (factorization and symmetry) were copied unchanged at the June checkpoint.new_hybrid_pdf.py, new_pdf.py, new_tsdadg_pdf.py, and LY93_pdf.py supplied distribution weights, means, and variances. One file was unchanged; three needed only import edits at that checkpoint.pdf_utilities.py supplied conversions between means, variances, and correlations. setup_clubb_pdf_params.py and hydromet_pdf_parameter_module.py contributed cloud/precipitation setup calculations; we extended their setup logic and interfaces.adg1_adg2_3d_luhar_pdf.py, pdf_closure_module.py, new_pdf_main.py, and new_hybrid_pdf_main.py were copied as starting points, then expanded to cover missing Fortran branches and connect our data objects and statistics. The original integrated path covered fewer options.stats_accumulate offered a reference for temperature, cloud-water, and vertical-average diagnostics. These were implemented using our existing JaxStats bridge and Fortran’s output order.precipitation_fraction.py, Nc_Ncn_eqns.py, and corr_varnce_module.py were consulted for precipitation coverage, droplet concentrations, and correlations. The amount of code carried over is less certain.How this was checked: archived dialogue and file actions, plus Git comparisons against CLUBB-JAX before the core replacement. Shared formulas alone do not establish copying: both ports follow the same Fortran.
The takeaway
F2PY kept the original program available. The same-input tests made every change measurable. Going depth-first forced one real branch through JAX early. Reusing those lessons made later translations bigger and less hands-on. By the end, “port this” could mean a nearly complete bundle of translated, tested, compared, commented, and formatted code—not just a draft function.
A runnable hybrid existed from the first replacement. Each piece could be reviewed and compared in a real run while the rest stayed in Fortran.
The full-run comparison supplied the main check throughout. Targeted routine tests were available for edge cases, not a separate test-writing project required for every file.
Differences could be traced while the changed piece was still small. Debugging still happened, but it could happen along the way instead of accumulating into one large final cleanup.
Compiled helpers, JAX patterns, and written instructions carried forward. Later tasks could focus on the next calculation and deliver code ready for human review.
The “miscellaneous related” bucket stays available below the stage overview, but it is not included in the headline totals. The report contains aggregate counts only; it does not publish prompt text.