Inside the magic that makes Odit audit
How 2,491 YAML files convert your bank SMS into structured data.

On this page
- What I'm actually parsing
- Why these are YAML files
- One pattern from start to finish
- The schema catches my typos
- Before any regex runs
- Identifiers have to stand on their own
- Faking a field that isn't there
- Codegen, and hashing everything into a version
- The page I actually live in
- The failure shapes
- What counts as a miss
- The health sweep
- What's left
Odit reads the your bank sends you after a transaction and turns them into an actual transaction list. There's no open banking API here, so those messages are the only data source we have.
The thing doing the reading is 2,491 YAML files across 21 providers, compiled down to 78,410 lines of TypeScript.
This post is what's in them, and mostly why. A good chunk of the "why" is just me getting it wrong the first time and having to go back.
Quick note before we start: the pattern file below is the real one, and the it reads is a real transfer of mine, twenty birr to a friend. As a rule the corpus doesn't leave the server, and a guard fails the deploy build when corpus text shows up in anything published. This file is the one exception I've carved out of it, because the made-up example I tried first kept flattening exactly the details worth showing.
What I'm actually parsing
21 banks and wallets, all of them a bit special, and all of them full of tiny variations that kept breaking my strict matching.
Telebirr alone has 501 pattern files. Not because telebirr does 501 different things, but because "you sent money" and "you sent money and paid a fee" and the Amharic version of each are separate strings with the fields sitting in different places.
Then there's everything in an inbox that isn't a transaction. OTP codes, promos, balance reminders. I have to be as good at throwing those away as I am at reading the real ones, because a promo misread as a transaction puts a fake number in someone's financial history.
Why these are YAML files
They used to be TypeScript. Moving them to YAML is the change that made this maintainable, for three reasons.
A pattern is data, and data wants a diff you can read. When a bank changes a template and I have to change a regex, the review should be the regex before and the regex after, on two lines. YAML is just funky formatted text, so that's exactly what the diff is: nothing else lives in the file to drag into it.
Tools that aren't the compiler can read data. Because a pattern is a file with a known shape, I've got a local admin app that loads it, runs a candidate regex against the real corpus, shows me what starts matching and what stops, and writes the file back. That loop is most of why the corpus is any good, and I don't get it when the definition is code.
Generated code can't be edited in place. YAML is the source of truth,
src/generated/ is output. If you hand-edit a generated file, the next codegen
run quietly eats it. Making that whole tree obviously machine-owned takes away
the temptation.
One pattern from start to finish
Say I send a friend twenty birr on telebirr. A few seconds later this lands. On a big enough screen it stays pinned while you read, and the parts being read light up.
Dear Robel You have transferred ETB 20.00 to dagim alemu (2519****5183) on 12/09/2026 00:08:40. Your transaction number is DIC6NJHLOC. The service fee is ETB 0.87 and 15% VAT on the service fee is ETB 0.13. Your current E-Money Account balance is ETB 6,127.70. To download your payment information please click this link: https://transactioninfo.ethiotelecom.et/receipt/DIC6NJHLOC. Thank you for using telebirr Ethio telecom
TRANSFER_TB_FROM_TB, the file that reads it, top to bottom.
name: TRANSFER_TB_FROM_TB # what every tool calls this pattern
provider: telebirr
category: transfer # money categories get the full health treatment; promos don't
order: 184 # tiebreak for when two patterns match the same body
status: done # triage state: draft, pending, flagged, done or junk
identifier:
pattern: '(?=[\s\S]*VAT)(?=[\s\S]*service fee is\s+ETB)You have transferred ETB [\d,.]+\s+(?!Tip\b)to [\s\S]+?Your transaction number is'
flags: isidentifier is separate from the field regexes. It decides whether this
pattern applies at all, and everything below only runs once it's said yes. The
two lookaheads at the front are doing real work: telebirr sends a nearly
identical template where fee and VAT are one combined line, and that one is a
different pattern. (?!Tip\b) keeps tips out of this one too.
regexes:
SENDER_NAME:
# The first capture group is the value that gets kept.
pattern: '^Dear\s*([\w\s]+?)\s*You '
flags: i
source: common/NAME_PATTERN
AMOUNT:
pattern: transferred ETB\s*([\d,]+(?:\.\d+)?)
flags: i
source: common/TRANSFERRED_ETB_PATTERN
TXN_ID:
pattern: transaction number is\s*([A-Z0-9]+)
flags: i
source: common/TBTXN_PATTERNFields are named, not positional. Everything downstream says TXN_ID, so
adding a field later doesn't renumber anything. And source means the regex
is a reference to a shared copy: 64 of them are common enough to live in one
file every pattern can point into.
RECIPIENT_NAME:
# Grabs everything up to the phone number or the date. The slot holds
# whatever the other side's account says, junk included, so a character
# class can't be written for it.
pattern: \bto\s+([^(]+?)(?:\s*[\(]|\s+(?:on|at)\s)
flags: i
RECIPIENT_PHONE:
pattern: (?:\(\s*(\+?[\d*]+)\s*\)|to\s+(\d{10,})\s)
flags: i
DATETIME:
pattern: (?:on|at)\s*([\d/:-]+\s[\d:]+)
flags: iThat loose capture on RECIPIENT_NAME is safe, because a name is never used
to identify an account downstream. Only a number is, and the phone lands in
its own field, mask and all.
SERVICE_FEE:
pattern: (?:The\s+)?service (?:fee|charge)(?:.*?)?is\s*ETB\s*([\d,]+(?:\.\d+)?)
flags: i
VAT_AMOUNT:
pattern: (?:\d+%\s+)?(?:VAT(?:\s+is\s+ETB|\s+on\s+the\s+service\s+fee\s+is\s+ETB))\s*([\d,]+\.?\d*)
flags: i
BALANCE:
pattern: balance is ETB *([0-9,]+(?:\.[0-9]+)?)
flags: i
RECEIPT_URL:
pattern: (https:\/\/transactioninfo.ethiotelecom.et\/receipt\/[^\s.]+)
flags: i
fallback:
template: https://transactioninfo.ethiotelecom.et/receipt/{TXN_ID}Fee and VAT capture separately and get summed downstream: this
cost me 20 birr plus exactly 1.00 in charges, 0.87 fee
and 0.13 tax. The fallback on RECEIPT_URL rebuilds the link from the
transaction number when a message doesn't carry one.
extraction:
direction: OUTGOING # from the account owner's side: money left
amounts:
principal: AMOUNT # the money that moved
fees:
- SERVICE_FEE
- VAT_AMOUNT
balance: BALANCE # what was left after
participants:
- fields:
name: SENDER_NAME
role: SENDER
type: CUSTOMER
- fields:
name: RECIPIENT_NAME
phone: RECIPIENT_PHONE
role: RECEIVER
type: CUSTOMER
transactionIds:
telebirr: TXN_ID
timestampField: DATETIMEextraction is where captured text stops being strings and becomes a
transaction. Participants carry roles: SENDER is me, RECEIVER is dagim.
That's what turns a list of phone numbers into a list of people.
The schema catches my typos
Every file gets validated on the way in:
export const PatternDefSchema = z.object({
name: z.string(),
provider: z.string(),
category: z.string(),
order: z.number().default(0),
status: z.enum(['draft', 'pending', 'flagged', 'done', 'junk']).default('draft'),
identifier: RegexDefSchema,
regexes: z.record(z.string(), RegexDefSchema).default({}),
extraction: ExtractionSchema,
});Misspell a role, point principal at a field that doesn't exist, name a debt
product that isn't in the registry, and the build stops. Otherwise you get a
pattern that runs fine and extracts nothing, which is a much worse afternoon.
If you make the wrong thing impossible to write down, you stop having to catch it later.
Before any regex runs
A goes through three layers, in this order.
Layer one is an ignored set. 95 exact bodies that get dropped without a word. Single characters, a bare account number with no sentence around it, truncated sends. It's a set lookup on the whole body so it costs nothing.
Each of these could have been a pattern instead. But there's nothing in them to extract, and they barely vary: a resend lands a character or two away from the last one. A pattern here is one more file in the 2,491 and one more identifier to keep from colliding, spent on a I never wanted to parse in the first place.
The set is silent though. A body that goes in disappears without a trace, so I only add one I'd be fine never showing a user again.
Layer two is an exact-body map. Promos go out in bulk and are byte-identical
to everyone. There's nothing in them to extract and no reason to run a regex, so
they sit in a Map keyed on the full body and a hit returns immediately.
This one exists because the thing it replaced was actively hurting me. Promos used to get caught by broad catch-all patterns, one per provider per language. Those matched on the shared sign-off text at the bottom, usually the bank's own name, which every transactional template also has. My collision counter hit 373,000 matching more than one pattern. That's about 78% of everything I supported at the time.
Deleting the catch-alls and matching those bodies exactly fixed it in one go. Nobody writes these entries by hand. A harvester Claude wrote walks the corpus instead: group by body, keep anything with 2+ duplicates, skip what a specific pattern already handles, write out YAML.
Layer three is the regex patterns. Whatever's left is templated transactional text with fields that move, and that's what the 2,491 files are for.
Identifiers have to stand on their own
The rule that keeps this from collapsing: an identifier has to be right on its own, without depending on which pattern got tried first.
Leaning on ordering works the day you write it and breaks the day someone inserts a pattern above yours. And it breaks silently. The gets filed as the wrong type, the fields still extract because the shapes are close enough, and nothing anywhere errors.
So identifiers use unique phrases plus negative lookaheads, the (?!Tip\b)
from earlier being one. A transfer pattern says the thing only a transfer
says, and explicitly says it isn't the tip or the reversal sharing
most of the same words.
There is still an order field, but it's a tiebreak, not a strategy. It's also
where a genuinely annoying bug lived. 143 patterns share order: 150, and my
sort had no second key, so when two tied the winner came down to whatever order
the filesystem handed the files back in, which could differ between my laptop and
the server.
The fix is one clause:
patterns.sort((a, b) => a.order - b.order || a.name.localeCompare(b.name));Two patterns matching the same body is still a bug, though. It gets counted, not
resolved by precedence, and the health sweep raises it as COLLIDES.
Faking a field that isn't there
Some fields aren't in the text but can be rebuilt from the ones that are.
You saw the real case above: telebirr's receipt URLs are just a fixed host with the transaction number stuck on the end. from before a 2025 format change don't carry the URL at all, but they do carry the transaction number, so the link can be rebuilt exactly.
fallback:
template: https://transactioninfo.ethiotelecom.et/receipt/{TXN_ID}It only fires when the regex missed and every field named in the template
matched. That second condition is the whole safety property. If a placeholder
doesn't resolve you get nothing, instead of a URL with the literal text
{TXN_ID} sitting in the middle of it pointing at a page that doesn't exist.
This template is also where a second product came from.
The page behind that URL is the provider's own statement of the transaction. My parse of the can be wrong; the provider's receipt page can't. And if the URL is just a host plus a transaction number, you don't need the SMS at all to get there.
That's useful to people who've never heard of Odit. A transfer screenshot is easy to fake and the page the bank serves is not, so when someone claims they've paid you, the receipt link settles it. It started at v.odit.et as a page you pasted a receipt link into; it outgrew the subdomain and it's links.et now, reading receipts from 18 providers.
Codegen, and hashing everything into a version
npx tsx src/codegen/index.ts reads every YAML file and writes
src/generated/. One patterns module and one exact-body module per provider,
the shared regexes, the ignored set, and a version.
The version is the fun part:
const hash = crypto.createHash('sha256');
// _common/regexes.yaml, then _ignored.yaml, then every _exact/*.yaml,
// then every provider directory, each file list sorted.
return hash.digest('hex').slice(0, 16);PATTERN_VERSION is a content hash over everything that can change a parse
result. Nobody bumps it by hand, so nobody can forget to.
Every directory read gets sorted before it's hashed, because readdirSync order depends on the filesystem, and an
unsorted hash would come out different on two machines holding identical files.
A version that changes when the content didn't is worse than having no version
at all, because it kicks off a full reparse of the corpus for nothing.
It also has to cover the ignored set and the exact-body maps, not just the regex patterns. I missed that at first. An earlier version of this hash only read the provider directories, so adding a promo body to the exact map produced no version change, nothing reparsed, and the it should have reclassified just sat there as unknown until something unrelated forced a run.
The page I actually live in
None of the above gets you a good corpus on its own. What does is that every pattern gets run against a few million real before I commit it.
The cheap version of that is just throwing a candidate regex at the table, which I can do because definitions are data:
select raw_body from sms_export_items where raw_body ~ $1 limit 200That tells you what a regex catches, and more usefully what it catches that you didn't mean. But it doesn't tell you whether the pattern's fields came out, and that's where the actual bugs are. A pattern can win on a and still leave the balance empty, and you'd never know from a match count.
So there's a page per pattern in the local admin app, and the top of it is a
whole-corpus extraction verdict. For TRANSFER_TB_FROM_TB it currently
reads:
SWEPT 100.0% of 115,034 extract cleanly
missing: VAT_AMOUNT 2 SENDER_NAME 2 BALANCE 1
Wins 115,034 messages from 664 people. Extracts fully on 99% of wins;
misses VAT_AMOUNT on 2, SENDER_NAME on 2, BALANCE on 1.
Derives RECEIPT_URL from TXN_ID on 2 wins.Five bad out of 115,034. I'd never have found those by paging through dots.
The failure shapes
The bit that saves the most time is that it doesn't list the five failures. It lists the shapes:
FAILURE SHAPES · ONE REAL MESSAGE EACH
2x SENDER_NAME
1x VAT_AMOUNT BALANCE
1x VAT_AMOUNTThree distinct ways this pattern can come up short, one real body shown for each. Fixing a shape fixes every in it, and the count tells me whether it's worth touching the regex at all. Two messages out of 115,034 usually isn't.
Next to it sits the richest fully-clean body as a reference, with every captured span highlighted in the itself, so you can see exactly what each regex grabbed rather than trusting that it did. On my transfer from earlier that's:
SENDER_NAME Robel
AMOUNT 20.00
RECIPIENT_NAME dagim alemu
RECIPIENT_PHONE 2519****5183
DATETIME 12/09/2026 00:08:40
TXN_ID DIC6NJHLOC
SERVICE_FEE 0.87
VAT_AMOUNT 0.13
BALANCE 6,127.70
RECEIPT_URL https://transactioninfo.ethiotelecom.et/receipt/DIC6NJHLOCSeeing the spans is how you catch a regex that matched the wrong thing. A lazy capture that runs one word too far still counts as a match, and it still looks green on any pass-fail check.
What counts as a miss
The verdict rule is in field-verdict.ts:
a message is green when every REQUIRED field either captures or renders from its fallback template, red when none do, yellow in between. Optional fields never count against a body; uncompilable regexes are invisible.
Required means !optional && !missVerified. RECEIPT_URL stays required here
because its fallback can always rebuild it. A field with no fallback that the
sender genuinely omits gets optional instead, or every old-format
sits here as a permanent miss and the three real shapes
drown under thousands of fake ones.
missVerified is the same idea with a human behind it: I looked, the sender
genuinely omits it, stop counting it against me.
And "Derives RECEIPT_URL from TXN_ID on 2 wins" is the fallback template doing its job in the wild. Those two don't carry a URL, the template rebuilt it from the transaction number, and they counted as clean rather than as misses.
Every field captures before anything is judged, because a required field's fallback template is allowed to reference an optional field's value. Judge as you go and the derive silently fails.
That same rule is used in three places, on purpose. The dots down the side of the
page, the API's failing-first scan, and the health sweep's accumulator all call
the same function, so they can't quietly disagree about what a broken is.
The page also shouts index predates current patterns when the sweep it's
reading is older than the YAML, because a green bar built from last week's
regexes is worse than no bar.
The whole app only runs on localhost and it's excluded from every deploy, because it's a tool that puts production regexes and real bodies on screen.
The health sweep
This is the piece that keeps a corpus this size from rotting without me noticing.
It walks the whole corpus once and prices every detector in that single pass, then flags patterns:
| Flag | What it means |
|---|---|
BAD-IDENT / BAD-REGEX | the regex doesn't compile |
COLLIDES | two identifiers match the same body |
DEAD | zero matches in the entire corpus |
SHADOWED | it matches, but another pattern usually wins |
UNPARSED | an amount field grabbed something non-numeric |
P-LOW / B-LOW | principal or balance under 98% coverage |
PARTIAL | over 10% of wins only extract some required fields |
RESIDUE | money-shaped tokens left uncovered on a lot of wins |
DIVERSE | one pattern matching loads of different body shapes |
AUTO-ROLE | a participant role never got triaged |
DEAD and DIVERSE are the two that keep earning their keep.
DEAD catches a pattern I wrote for a format the bank has since dropped. It
compiles fine and it looks maintained, so nothing about the file tells you it
stopped matching anything months ago. You only find out by counting across the
whole corpus.
DIVERSE catches the opposite: an identifier loose enough to match free text,
winning across a pile of differently-shaped bodies. That's a pattern turning into
a catch-all, which is the exact thing the exact-body layer was built to undo. Nice
to catch it before the collision count blows up again.
Two flags can be silenced per pattern, and both need me to say so in the file.
residueVerified when I've looked at the leftover tokens and they're carrying
nothing, and missVerified when a field is genuinely missing upstream rather
than being missed by me. Measurement keeps going either way, only the flag stops.
The output gets compared against a committed baseline, so making things worse shows up as a diff instead of as a support six weeks later. Accepting a new baseline is a commit, which makes it a decision somebody made instead of drift.
What's left
- Content-hash identity end to end, so the row id stops being load-bearing
- Auto-proposing YAML for unmatched clusters instead of me writing each one
- A proper story for a provider changing a template mid-flight, rather than
waiting for
DEADto tell me weeks later - The Amharic patterns are thinner than the English ones and I know it
The user side of all this is that they don't do anything. Install the app, grant access once, and the history builds itself out of what's already on the phone, including everything from before they installed it. They never download a statement and they never hand a password to anybody.