CLUBB → JAX · focused effort story

How the port learned
to port itself

The short version: first we built a bridge to the old Fortran program. Then we built a measuring stick. Then we finished one difficult path through JAX. After that, the work changed from figuring out how to port CLUBB into applying a recipe that already worked.

Code added, with and without tests ↗ · Tokens by stage ↗

00

Thirty-second orientation

CLUBB is an atmospheric model. Its trusted implementation is written in Fortran.

JAX is a Python numerical system that can compile calculations for CPUs and accelerators.

Porting here means recreating a Fortran calculation in JAX while keeping its numerical behavior checkable.

A human touch is one human-written prompt during an agent task.

01

The recipe

Six stages of the port

Pick a stage to see what it was for, what it produced, and what the human/agent rhythm looked like. These are stages of the method, not six strictly separate time periods.

01b

Watch the method

A working port.
From the first file.

Start at the top, then follow one branch all the way down. Check the whole run after each piece. When you move sideways, shared helpers are already JAX—ready to use again.

Build JAX one branch at a time, validating a complete run after each replacement Fortran stays on the left while JAX grows on the right. Core has four branches. Finish Mixing all the way down to a shared Grid helper, then move to Clouds, Transport, and Pressure. Blue links show those branches reusing the same JAX helper, not creating copies. Dashed amber arrows borrow unported Fortran routines through F2PY. A scanner compares full runs after every new piece. At the end no temporary Fortran calls remain.
Save GIF ↗

Illustrated call tree; each box stands for a piece of code, not a literal file or the complete CLUBB call graph. Swipe across the diagram to see both trees.

01 / Always runnable

Borrow the pieces you haven’t ported yet.

The first JAX entry point already runs a case. It calls the remaining Fortran routines through F2PY—the bridge that lets Python call Fortran. The port works from the start, even while it is “cheating.”

02 / Check the whole run

No separate unit test needed for each piece.

Run the hybrid program and the original Fortran program with the same inputs after each replacement. The selected CLUBB cases check real callers, shared state, and outputs together. A new unit test for every routine is not a prerequisite: the full-run comparison is already there. Small unit tests can still stress-test edge cases the selected runs do not exercise.

03 / Go deep, then reuse

Port a helper once. Use it in many branches.

Depth-first means finishing one call path all the way down before moving sideways. Different Fortran branches often use the same small routines. Once those helpers are ported, compiled, and checked, later branches can call their existing JAX versions. We reuse working code, not just lessons learned—and keep each change small enough to review.

01c

The code and the checks around it

How many lines were added?

We count lines added in Git commits connected to the port. One view includes tests; the other leaves them out. Both include support scripts, configuration, comments, and blank lines.

Download all stage counts ↗

Added lines are not final program size. Rewriting a line can count it again in a later commit. Formatting changes count too. Code that was never committed is outside this measure.

What does this count include?

C1 includes tests; C2 excludes them. The chart selector applies to every code chart on this page. Documentation is counted separately, not included in either measure.

02 / 03

How people talked to the agents

How long were the prompts?

Distribution of human-written prompt lengths Bars show how many reviewed prompts fall into each word-count range.

Some short messages were simply “go” prompts—such as “continue,” “go ahead,” or “OK, go for it, do a good job”—rather than new instructions. A short prompt can still contain an important correction or decision.

What the human touches did

Build, steer, ask, repeat

Length is one view; purpose is another. Each row is a stage, and each colored width is the share of its prompts assigned to that role.

Each prompt gets one category using wording-based rules. A message can do several things, so these labels are a guide to the conversation—not a precise judgment of intent or necessity.

04

The part we care about most

The core port became a repeatable job

The first complete dependency path taught the rules. The next wave reused them across much more code.

A ready-made check

The same-input full-run tests were already established. Each new branch could use that comparison, with additional cases and checks added as problems surfaced.

A written recipe

Markdown instructions and reusable agent commands captured how to port, run comparisons, and finish a change. Later tasks could refer to that workflow instead of restating every step.

Problems solved at a small scale

JAX-compatible objects, array shapes, branching, and updates had been worked through on one path. Those patterns carried across branches, and already-ported leaf routines could be called again directly.

This is the practical change: early prompts often named specific routines, edge cases, and JAX constraints. Later tasks could hand over a much larger block and ask for the full package: port it, test it, format it, and check it against Fortran.
05

Real calendar time

How many committed lines landed per human prompt?

Each bubble is one stage, placed at the middle of its prompt activity. Bubble size shows prompt count.

Committed line additions per human-written prompt over calendar time Six stages are positioned by their activity midpoint and measured by committed line additions per prompt.

This shows how the size of committed changes per human touch evolved as the code and porting workflow matured. It is not a model leaderboard: stages had different scopes, and the chart does not control for model differences or reasoning settings. Formatting and repeated rewrites also count as added lines. The first core path and later core scale-out are the closest comparison here.

06

Inside each stage

Where did the added code go?

The stages tell us what the work was for. These bars break down the code it changed: calculations, the driver around them, tests, and supporting tools.

Hover or tap a colored segment for its count.
Committed line additions by component across all six stages Six stacked bars show the code components changed in each stage, as shares or on a common added-line scale.

What are these parts of the code?
Core calculations
The central CLUBB calculations and their numerical helpers.
Driver
The code that sets up a case, advances the run, and connects its parts.
Statistics
Diagnostic calculations and output used to inspect a run.
Radiation
Calculations involving energy transferred by sunlight and thermal radiation.
Microphysics
Cloud and precipitation processes, such as changes involving water droplets or ice.
Build + runtime tools
Support for building, configuring, launching, and connecting the program.
Tests
Comparison runners, validation helpers, and other test code—not only unit tests.

Categories follow the files changed, not the reason for every individual line. One stage can touch many parts of the program. Documentation stays outside these code counts.

06b

The work behind the text

Tokens used at each stage

Tokens are pieces of text processed or generated by a model—not lines of code. Input includes instructions, conversation history, code, and tool results, often read again on later calls. Output includes explanations, tool commands, code, and reasoning.

Download stage + model breakdown ↗

T1 = T2 + T4. Cached input is already in T2; reasoning is already in T4. Do not add them again. Large input counts include repeated context; they are not billions of newly written words.

Input, cached input, output, reasoning—what is the difference?
T2 · Input: what the model reads
Instructions, earlier messages, code, and tool results supplied on a model call. The same context can be counted again on later calls; this is not a count of unique text.
T3 · Cached input: reused reading work
When the beginning of a request matches previously processed context, the system can reuse that processing. These tokens are still input, so T3 is already inside T2. Caching does not mean the answer itself was reused.
T4 · Output: what the model generates
New text and tool calls, including code, explanations, and any recorded reasoning tokens. It is broader than the final answer a person sees.
T5 · Reasoning: part of generation
Tokens used for the model's internal reasoning before or while producing its response. The usage record gives a count, not the reasoning text. T5 is already inside T4.
T1 · Total: input plus output
Add T2 and T4 once. Cached input and reasoning describe portions of those totals, not extra tokens to add on top.

A tiny example: 1,000 input tokens (including 800 cached) + 200 output tokens (including 50 reasoning) = 1,200 total. The subsets show how the tokens were used; they do not increase that total.

Token counts are not a subscription invoice or a count of code lines. Official token definitions ↗

06c

Building on earlier work

CLUBB-JAX helped the core scale-out.

Alex Connolly’s CLUBB-JAX gave us reusable numerical routines and a reference for other translations. Its clearest contribution came as we expanded the core beyond the first completed branch.

03 · First core path

Examples for faster JAX calculations

Comparisons with advance_xm_wpxp and advance_xp2_xpyp informed updates across whole atmospheric columns. We also copied parameter_indices.py, adapting its name-to-array-position lookup to Python’s zero-based indexing.

04 · Core scale-out

Reusable cloud-distribution mathematics

We copied 13 PDF-related files. PDF means probability distribution: variations in moisture, temperature, and motion within a grid cell. Some helpers carried over almost directly; larger routines needed expansion and integration.

The whole core, file by file

56 source files at the June 16 core checkpoint: 49 core files plus seven supporting data-object files. Package markers, tests, the driver, and later microphysics are excluded.

Each file counts equally. Categories describe the June code and involve review judgment; they do not measure lines borrowed or hours saved. Percentages are rounded.

Explore all 56 files and their categories

Each note distinguishes direct evidence, a review judgment, or an unresolved case.

    What carried forward

    Useful building blocks, with a broader port built around them.

    Root solvers, matrix helpers, and many cloud-distribution formulas carried forward; the root and matrix calculation bodies still match in September. Around them, we expanded missing branches, connected statistics and data objects, and validated complete runs against Fortran.

    What carried over, and what we built around it
    • Root and matrix calculations: calc_roots.py (polynomial equations) and matrix_operations.py (factorization and symmetry) were copied unchanged at the June checkpoint.
    • Cloud-distribution formulas: new_hybrid_pdf.py, new_pdf.py, new_tsdadg_pdf.py, and LY93_pdf.py supplied distribution weights, means, and variances. One file was unchanged; three needed only import edits at that checkpoint.
    • Parameter setup: pdf_utilities.py supplied conversions between means, variances, and correlations. setup_clubb_pdf_params.py and hydromet_pdf_parameter_module.py contributed cloud/precipitation setup calculations; we extended their setup logic and interfaces.
    • Coordinating the calculations: adg1_adg2_3d_luhar_pdf.py, pdf_closure_module.py, new_pdf_main.py, and new_hybrid_pdf_main.py were copied as starting points, then expanded to cover missing Fortran branches and connect our data objects and statistics. The original integrated path covered fewer options.
    • Statistics: the original stats_accumulate offered a reference for temperature, cloud-water, and vertical-average diagnostics. These were implemented using our existing JaxStats bridge and Fortran’s output order.
    • Further reference material: precipitation_fraction.py, Nc_Ncn_eqns.py, and corr_varnce_module.py were consulted for precipitation coverage, droplet concentrations, and correlations. The amount of code carried over is less certain.

    How this was checked: archived dialogue and file actions, plus Git comparisons against CLUBB-JAX before the core replacement. Shared formulas alone do not establish copying: both ports follow the same Fortran.

    07

    The takeaway

    The important output was not only JAX code. It was a porting method.

    F2PY kept the original program available. The same-input tests made every change measurable. Going depth-first forced one real branch through JAX early. Reusing those lessons made later translations bigger and less hands-on. By the end, “port this” could mean a nearly complete bundle of translated, tested, compared, commented, and formatted code—not just a draft function.

    No wait for a whole rewrite

    A runnable hybrid existed from the first replacement. Each piece could be reviewed and compared in a real run while the rest stayed in Fortran.

    No unit-test suite required before porting

    The full-run comparison supplied the main check throughout. Targeted routine tests were available for edge cases, not a separate test-writing project required for every file.

    No need to save integration for the end

    Differences could be traced while the changed piece was still small. Debugging still happened, but it could happen along the way instead of accumulating into one large final cleanup.

    No need to solve every branch from scratch

    Compiled helpers, JAX patterns, and written instructions carried forward. Later tasks could focus on the next calculation and deliver code ready for human review.

    Scope and counting notes

    The “miscellaneous related” bucket stays available below the stage overview, but it is not included in the headline totals. The report contains aggregate counts only; it does not publish prompt text.