Kolen Cheung
  • Posts
  • Graph
  • About
Categories
All (37)
Agentic engineering (16)
Astronomy (1)
Cultural differences (10)
Cybersecurity (1)
Durham HPC Days (2)
Formal methods (1)
JAX (3)
JIT (1)
Julia (1)
JWST (2)
Link (10)
LLM (19)
Machine learning (2)
Math (18)
Numba (2)
Open source (1)
Pandoc (1)
Physics (2)
PyCon UK (1)
Quantum mechanics (1)
References (1)
Reproducibility (4)
Research ethics (13)
Research software engineering (14)
RSECon (1)
Statistics (3)
Talk (9)
Theory building (4)

Posts

The economy of AI math, for mathematicians

Math
LLM
Research ethics
Cultural differences

Why would OpenAI keep testing open problems on its internal models, when the mathematicians’ advisory group asked it to stop? My guess: because of how an LLM is paid for.

October 9th, 2026

The flood, released responsibly?

Math
LLM
Research ethics
Cultural differences

OpenAI released 722 AI-generated manuscripts after consulting the mathematicians’ advisory group. An improvement on the Navier–Stokes incident, short of what the group asked, and the flood Gowers predicted.

October 8th, 2026

What I found interesting at RSECon26

Talk
RSECon
Research software engineering
Agentic engineering

A five-minute report of the talks I found interesting at RSECon26: AI and the value of RSEs, the human work of research software, and how other groups are run.

October 7th, 2026

A proof with no warranty

Link
Math
Open source
Cultural differences

Thomas Bloom asks mathematicians not to abandon mathematics as a human activity. For some open source maintainers, software never was one, which is why OpenAI dropping a Navier–Stokes proof looks so different from each side.

October 6th, 2026

What do research software engineers uphold?

Talk
Agentic engineering
Research software engineering

My application to the Software Sustainability Institute Fellowship 2027: getting research software engineers to talk about AI, and to say what they value.

October 5th, 2026

Mathathon, redesigned

Link
Math
LLM
Research ethics

The Mathathon organizers and two coauthors of the open letter against it redesigned the event: old problems, new proofs, and no sponsorship from proprietary AI labs.

October 4th, 2026

What does understanding mean in astronomy?

Link
LLM
Astronomy
Research ethics
Research software engineering
Cultural differences

Jess Werk, chair of Astronomy at the University of Washington, on protecting the Ph.D. from AI — and how differently astronomy, CMB science and RSE might mean “understanding”.

October 3rd, 2026

Extracting metadata with an LLM is data cleaning

LLM
Statistics
Research software engineering
Reproducibility

Your LLM metadata extraction gives different answers each time you run it. I think that is the uncertainty showing, and the way to deal with it is the data cleaning we did before LLMs: a statistical problem, with a human checking a small sample.

October 2nd, 2026

Should I use an LLM to refactor my legacy code?

Agentic engineering
Research software engineering

You have a legacy codebase and think maybe an LLM could help refactor it. Here is the flowchart I’d walk through to decide, and why it is in that order.

October 1st, 2026

Mathematics is the canary

Talk
Math
Agentic engineering
Research software engineering
Cultural differences

Mathematicians now have what software engineers can only wish for: a complete specification and a machine-checked proof. What happened to them when agents got there, and what research software engineers should take from it.

September 30th, 2026

Gowers chose good old-fashioned AI in 2022

Link
Math
LLM

Tracing back to Timothy Gowers’ 2022 job opening for an automatic theorem proving project, and why he wanted AI that comes with understanding.

September 29th, 2026

Can a theory be laid out, and can it be rewarded?

Agentic engineering
Theory building
Machine learning

Can understanding be rewarded, the way verifiable signals reward correctness? I think it depends on whether a theory can be laid out first, and I only have an answer to that for science.

September 28th, 2026

Where do correctness and understanding live in research software?

Research software engineering
Agentic engineering
Theory building

Charity Majors says truth lives in the tests and we will ship code we have never read. Peter Naur says a program’s theory can’t be revived from its documentation. For scientific computing, I’m looking for somewhere in between.

September 27th, 2026

MCMC for a logistic fit from early data

Statistics

Following up the Fisher information post: simulate early data from a logistic, sample the posterior with MCMC, and see how little the data say about the limiting value.

September 26th, 2026

Fisher information for a logistic fit from early data

Link
Statistics

John D. Cook on why a logistic fit from early data can’t predict its limit — and the Fisher information that says how bad it is before fitting anything.

September 25th, 2026

Episteme and techne in AI for science

Link
Machine learning
Physics
Research software engineering
Cultural differences

Tapio Schneider on understanding and trust in AI for climate models and CFD: the same words as the AI-in-mathematics debate, meaning something else.

September 24th, 2026

Where to have your say on math and AI

Math
LLM
Research ethics

A short list of the high-profile declarations and letters on AI in mathematics, who can sign or endorse each one, and what you are agreeing to when you do.

September 23rd, 2026

Releasing pandoc-amsthm v3

Pandoc
Agentic engineering

pandoc-amsthm, the pandoc filter giving any output format the theorem environments of LaTeX’s amsthm, is reimplemented as a Lua filter, agentically.

September 22nd, 2026

What if the button came with the understanding?

Math
LLM
Research ethics

A thought experiment: suppose an AI lab could land a theorem together with the understanding. Would dropping it on mathematicians still be wrong?

September 21st, 2026

Suspending the Mathathon won’t stop the flood

Math
LLM
Research ethics

A coalition of current and former Caltech mathematicians has asked for the Mathathon to be suspended. They have quite a few points, but the economics are wrong: the $20,000 of AI credits they say no mathematician could reasonably expect is about what an individual subscription already buys over several months. And the letter assumes that if you don’t welcome AI labs to attempt mathematics, the flood will stop.

September 20th, 2026

A flood from within the community

Link
Math
LLM
Research ethics

Timothy Gowers on why he didn’t sign the Fields medallists’ declaration: a flood of AI proofs we can’t digest may still be a pretty good bargain.

September 19th, 2026

Installing things on someone else’s computer

Agentic engineering
LLM
Research software engineering
Reproducibility

A git pre-commit hook, uvx or npx, and AGENTS.md all install something on someone else’s computer. So when are you making these decisions?

September 18th, 2026

The technical debt of AI-generated mathematics

Link
Math
LLM
Agentic engineering
Research ethics
Research software engineering
Cultural differences

Henry Cohn on math slop as technical debt — and why the people left to pay it sound a lot like open source maintainers.

September 17th, 2026

Long-horizon agents write dead programs

Math
LLM
Agentic engineering
Theory building

Long-horizon tasks seeking a verifiable reward give you something logically sound yet indigestible to a human. In Peter Naur’s terms, it is a dead program.

September 16th, 2026

Deep theorems were scarce, and AI has broken that signal

Link
Math
LLM
Research ethics
Research software engineering
Cultural differences

Bryna Kra on realigning mathematics’ incentives once AI makes theorems abundant — and why that reads like the RSE credit problem.

September 15th, 2026

Why I do mathematical research

Link
Math
Theory building

Emily Riehl on theory building: finding the “right” level of generality so more people can hold complicated thoughts in their heads.

September 15th, 2026

What RSEs should take from the OpenAI Navier–Stokes incident

Math
LLM
Agentic engineering
Research software engineering
Research ethics
In the week of 8 September, OpenAI announced a proof of the forced blowup case of the Navier–Stokes Millennium Prize Problem. I have written a trio of blog posts surrounding…
September 14th, 2026

Mathematicians aren’t mourning their craft

Math
LLM
Agentic engineering
Research software engineering
Reproducibility
Cultural differences

Software engineers have learnt to read the reaction to coding agents as a split between people who love the craft and people who care about the product. Seen through that lens, mathematicians upset about OpenAI’s Navier–Stokes result look like craft-lovers in denial, moving the goalposts. I think that misreads them. “After Math”, by Silvia De Toffoli and Eamon Duede, argues that mathematics was never after the answer: the understanding is the product. Then I extrapolate to research software, where I’d argue the same holds for code that establishes scientific results.

September 13th, 2026

Progressive disclosure for mathematical proofs

Math
LLM
Agentic engineering
Research ethics
Cybersecurity
Cultural differences

The companion article covers the Navier–Stokes “incident” itself. Here I want to look at a parallel: AI agents are finding security bugs and producing mathematical proofs faster than the people working in these fields can absorb them.

Can mathematics borrow from coordinated vulnerability disclosure? I’d propose a registry that records work as it develops, with timestamps that others can verify, together with circles that gradually widen before public release. This would give people a record of their contribution without having to rush an unfinished proof out. But it only helps if we also give credit to ideas and partial results.

September 12th, 2026

Unpacking the Navier–Stokes blowup: the mathematics, the machines, and the mathematicians

Math
Physics
LLM
Agentic engineering
Formal methods
Research ethics

OpenAI claims a solution to alternatives C/D of the Navier–Stokes Millennium Prize Problem. In the same week, two other groups announced finite-time blowup results for Euler using different approaches. This post is trying to unpack these results for someone with undergraduate mathematics and physics and some knowledge of LLMs and agentic engineering.

We start from the physics, write down the four Clay statements, and go through the idea of the construction. A collapsing vortex spins up by conservation of angular momentum, the same reason water turns faster as it nears a drain. High-frequency corrections provide the averaged momentum flux needed to keep the force smooth. This also explains how the speed can diverge while the total energy stays bounded, why Navier–Stokes blowup requires unbounded speed, and how rescaling gives the construction for every positive viscosity.

Then we discuss Lean, which checked two of the three results before human review. A proof can be correct and still be hard to learn from, as illustrated by an author calling his own Lean-verified paper “AI slop”. We still need to check that the formal statement says what we intend and that the proof uses only the allowed assumptions. On the ethics, I think OpenAI probably did not take anyone’s work; the different approaches support that. But two groups rushed unfinished work out, and Tao worries that open problems are being mined non-renewably. I think this is related to how mathematics assigns credit, from Newton–Leibniz to Perelman to the reception of computer-assisted proofs. Could open development help? We also discuss what happens when proofs cost millions of dollars, and some results may be worth keeping secret. Finally, we look at the role of rumours, scale, and verifiable targets, why I doubt mathematics will have an AlphaGo moment, and what the mathematical community and AI labs could do differently.

September 11th, 2026

A Primer on Just-In-Time (JIT) Compilation

Talk
JIT
JAX
Numba
Julia
Agentic engineering

Numba, JAX and Julia are the three just-in-time compilers most likely to be found in production scientific computing. Run the same small problems through all three and what separates them turns out not to be speed: each language design comes bundled with a different set of abstraction layers — its own tower of intermediate representations, its own introspection tools, its own notion of what you are allowed to say — and the optimisation strategy is baked into that bundle rather than chosen independently of it. The useful question about any of them is never “is it fast?” but which language am I actually writing? — and the answer is never the one in the file extension. Every speed-up below is a rewrite that computes exactly the same thing: a compiler pass performed by hand. What the three differ in is how much of that narrowing they leave to you.

Each system is walked down its own tower, with a runnable notebook behind every figure: why fusing four array allocations into two is where the speed actually comes from; why naming an intermediate costs Numba a 127 MB round trip that JAX’s tracer never pays, and why hand-writing the loop then beats array notation; why Julia’s fusion is syntax you write rather than an inference the compiler makes, why attaching physical units can cost exactly nothing, how a routine ustrip on a symmetric sparse matrix silently densified a matrix that would have been 450 TB, and what it means that a language is its own metaprogramming language. Then what the layering itself buys — separation of concerns, a place to verify, a place to optimise, a place to express intent — what compiling late adds on top of it, and what it costs: no artefact to archive, no code to sign, verifiability traded for performance.

Then a digression — an LLM read as a JIT compiler: the ultimate late binding and the ultimate metaprogramming language, but performing enormous narrowing with no specification licensing it and no guard checking it. Bun’s 535,000-line Zig-to-Rust rewrite worked because it had both.

The slides, this write-up and the PDF are generated from a single source, so everything that was on screen is here, together with what was said around it.

August 5th, 2026

Coding With Agents: What’s Different, What’s at Stake

Talk
LLM
Agentic engineering

Journal club discussion on agentic coding: the shift from autocomplete to autonomous agents (Claude Code, Codex CLI), vibe coding vs. agentic engineering, and the implications for software quality.

May 6th, 2026

Reproducibility in Computational Research

Talk
Research software engineering
Reproducibility

Reproducibility is a cornerstone of computational research, yet achieving it across diverse computing environments remains challenging. This examines the landscape of software environment management tools and their role in enabling reproducible research workflows. We begin by clarifying the terminology around reproducibility, distinguishing between re-runnability, repeatability, reproducibility, reusability, and replicability. We then explore the multifaceted nature of package managers—as software tools, ecosystems, indices, and distributions—and categorize them by scope, distribution method, platform support, and linking strategy.

The core focus is on practical solutions to three critical problems: reproducing research software environments across different systems and platforms, customizing build processes while maintaining reproducibility (particularly for high-performance computing contexts), and distributing these environments effectively. We provide in-depth analysis of several key technologies: Conda/Mamba for cross-platform binary package management, (TODO: Pixi for modern Python project workflows with lock files and task automation, and brief overviews of Nix, Spack, and Docker). Through concrete examples ranging from pure Python packages to complex scientific software stacks and HPC system environments, we demonstrate how these tools address real-world reproducibility challenges. The presentation emphasizes the importance of understanding the ecosystem around package managers, including channels like conda-forge, and discusses advanced topics such as build customization, platform-specific optimizations, (TODO: and the trade-offs between source and binary distributions).

November 26th, 2025

JIT compilers for scientific computing in Python: Numba vs. JAX

Talk
JAX
Numba
JWST
PyCon UK

Accelerating scientific Python with JITs. We share our journey migrating a gravitational lensing likelihood calculation from Numba to JAX. Learn about performance gains, automatic differentiation benefits, and practical lessons for high-performance scientific computing in Python.

Python is widely used in scientific research, but pure Python can sometimes be too slow for computationally intensive tasks. Just-In-Time (JIT) compilers are essential tools for boosting performance, allowing Python code to run closer to native speeds. While libraries like Numba have long been popular for accelerating numerical Python functions, JAX offers a new paradigm, combining JIT compilation with powerful features like automatic differentiation (auto-diff) and execution across different hardware (CPUs, GPUs, TPUs).

This talk will take you on a journey through our experience optimizing a critical component of an astrophysics analysis pipeline: the calculation of the likelihood function for gravitational lensing models, used with data from the James Webb Space Telescope. We initially used Numba to accelerate this calculation, but the need for performance portability across hardwares, and the potential speed up from gradient computation for model fitting led us to explore JAX’s unique capabilities.

This session will walk through the practical steps, challenges, and insights gained from migrating this complex scientific code from its existing Numba implementation to a JAX-based one.

You will learn:

  • Why leveraging performance tools like JITs is crucial for cutting-edge scientific analysis in Python.
  • The practical considerations when migrating existing numerical code from Numba to JAX, including syntax changes and managing state.
  • How JAX’s auto-differentiation simplifies gradient calculations essential for scientific optimization and sampling tasks.
  • The significant performance improvements achieved in our specific gravitational lensing case study by using JAX’s compiled functions.
  • Broader lessons learned about structuring scientific Python projects to effectively use modern JIT compilers and harness capabilities like auto-diff.

We’ll conclude by comparing Numba and JAX in benchmark performance, developer ergonomics, and tradeoffs between the two, providing you with practical guidance for choosing the right tool for your scientific computing needs.

This case study offers a concrete example of how evolving Python libraries are enabling researchers to perform complex, high-performance computations directly within the Python ecosystem. Join us to see how tools like JAX are empowering scientific discoveries, one optimised function at a time.

This talk is suitable for intermediate Python programmers familiar with NumPy.

September 21st, 2025

Durham HPC Days 2025

Talk
Durham HPC Days

Conference recap of Durham HPC Days 2025, covering themes of GPU acceleration, AI/HPC convergence, benchmarking, and the RSE community.

June 7th, 2025

Scientific Computing with JAX

Talk
JAX
Durham HPC Days
JWST

Presented at Durham HPC Days 2025: adapting a gravitational lensing likelihood from Numba to JAX, for James Webb Space Telescope observations.

June 4th, 2025

Quantum Mechanics references

References
Quantum mechanics

Annotated bibliography of quantum mechanics textbooks and lecture notes, with recommendations for different starting points and learning goals.

April 11th, 2025
No matching items

    © 2025–2026 Kolen Cheung