This one has an idea in it that I can't find anywhere else. That's the interesting part, and also the part most likely to be wrong.
Get hydrogen into a piece of high-strength steel and it turns brittle. Not immediately — that's the strange bit. You can load the part up, and it holds. Hours later, sometimes days later, with nothing having changed, it snaps. It's been known for 150 years. In 2013 it snapped 32 of the 96 big anchor rods on the new Bay Bridge in San Francisco within two weeks of their being tightened, months before the bridge opened. In 2014 it broke bolts on a London skyscraper, and two of them fell to the street. And it's one of the main reasons you can't simply pump hydrogen through the natural gas pipelines already in the ground.
Two explanations have been fighting about it for about fifty years.
One says hydrogen greases the works. Metals bend because tiny defects in the crystal slide past each other. Hydrogen makes that sliding easier, so all the deformation piles into one small spot instead of spreading out, and that spot tears.
The other says hydrogen weakens the glue. Hydrogen collects at the boundaries between metal grains, and weakens the atomic bonds holding those surfaces together, so they simply pull apart.
Both are plausible. Both have evidence. Fifty years in, neither has knocked the other out. The usual answer is that both happen, and nobody can say how much of each.
Why it's stuck
Here's what I think is the interesting part of the impasse. Most of the effort to settle it has gone into trying to see it. Put the sample under a better microscope. Get it into a synchrotron. Image the crack tip while it's growing and watch what the metal does right at the point of failure.
That's a reasonable instinct and it has produced fifty years of beautiful work. It's also the specific measurement that has been attempted most, and the two theories are still both standing after it. There's a reason: the action happens at a boundary buried deep inside a loaded piece of metal, and most of our best instruments need a clean, exposed surface to look at. You end up photographing somewhere the mechanism isn't.
Stop watching the metal. Watch the clock.
The idea I want to describe doesn't look at the steel at all. It looks at when things break.
Take a lot of identical samples. Load them all the same way, with the same amount of hydrogen in them, and wait. Some break in an hour, some in a day, some in a week. Write down the times. Now look at the shape of that spread.
The two theories predict different shapes, and they predict them for reasons that have nothing to do with metallurgy.
If hydrogen is greasing the works, damage builds up gradually. Small amounts of deformation pile up, a bit at random, until they cross a line. The signature of that kind of process is in how the risk changes over time: a bar's chance of breaking in the next hour climbs at first, then settles to a steady level. A bar that has hung there a long time is in no more danger than one that has hung there a medium time.
If hydrogen is weakening the glue, nothing can happen until hydrogen has had time to collect at some grain boundary past a critical level. So there's a quiet early stretch where almost nothing fails. After that it's a weakest-link situation — the part fails when its worst boundary gives, and a part has an awful lot of boundaries — and the chance of breaking in the next hour climbs steeply and keeps climbing. The longer it's been hanging there, the likelier it is to go.
You don't need a new machine to do this. You need a lot of samples and a stopwatch, and the patience to run the test properly.
The data mostly already exists
This is the part that made me sit up. Industry already runs a version of this test. If you sell bolts or pipeline steel or anything else that has to survive in hydrogen, there are standard qualification procedures where you hang a load on a batch of samples and see which ones break. People have been doing it for decades. The data is sitting in files.
And there's a quirk in how that data gets used. Tests have a deadline. When time's up, some samples have broken and some haven't. The ones that haven't are called run-outs, and they're usually recorded as a pass and set aside.
But a sample that survived 1,000 hours is not a blank. It's a fact: this one lasted at least a thousand hours. Statisticians call that censored data, and there's a whole mature field — the one insurers and drug trials run on — built specifically to squeeze information out of it. And the run-outs matter here more than usual, because the two curves above agree at the start and disagree at the end. The late end is exactly the part the deadline cuts off.
Where this came from, and why I think it's worth saying out loud
I didn't come up with this. I found it using a tool we build called blueMonster, which generates a large pile of candidate ideas for a question and sorts them. It produced 274 ideas for this one. Around 95% of them were the same move: pick an instrument, claim theory A gives one reading and theory B gives another. More microscopes.
This one was sitting at about 271st of 274, never developed any further. I only saw it because I read to the bottom of the list.
There's a reason it looks so different from the rest, and I like it enough to mention: alongside the metallurgy, I'd fed the tool a pile of concepts from credit risk. And this idea is, structurally, a borrowed argument from finance. There are two competing pictures of how a company defaults — one where its value drifts downward until it crosses a line, and one where it's fine until it suddenly isn't. Those two pictures produce different distributions of when defaults happen, and separating them from timing data is a normal thing to do in that field. Swap company for steel bar and default for fracture and you have the idea above.
Then I went looking to see whether anyone in metallurgy had already done it. There's plenty of nearby work — people use statistics on these tests all the time, to predict which materials are more susceptible, to rank alloys, to set safe limits. What I couldn't find was anyone using the shape of the failure-time distribution to argue about which mechanism is operating.
The honest problems with it
Three, and the first is the biggest.
It needs a lot of samples, and the two shapes are close cousins. They agree early and only pull apart late. Telling two curves like these apart with 90% confidence takes something like 60 exact failure times per condition, with no run-outs at all. A standard qualification test uses three or four samples for 200 or 720 hours, at a load chosen so that most of them survive. So no single test file settles anything. You'd have to pool many batches, and pooled batches bring their own mess — different heats, different platings, different amounts of hydrogen. That's why the existing data is the asset, and also why it's the problem.
The idea as I found it was thin. It said to analyse the distribution of failure times to tell the theories apart. It did not say which theory predicts which shape. The predictions I described above are a reconstruction — the reasoning is standard, but somebody has to do that derivation properly and commit to the direction before looking at data, or this is just curve-fitting with a story attached.
Reality is probably mixed. Most people in the field think both mechanisms operate together, which would give a blend of the two shapes rather than a clean one or the other. That sounds like a problem and I think it's actually the best part: fitting a blend and asking how much of each gives you a number, which is more than the yes/no argument has managed in fifty years.
A note on how to read this: I'm not a metallurgist and I have run no experiment. I've described an idea and told you how hard I looked for prior work, which was: moderately. If you work on hydrogen embrittlement and this is old news, please tell me and I'll say so here. If it isn't, it's yours — I'm not doing anything with it, and I'd be delighted to read the paper.