Explainer··16 min read

Inside the magic that makes Odit audit

How 2,491 YAML files convert your bank SMS into structured data.

A wall of patterned ceramic tiles, every square carrying a different design
On this page

Odit reads the SMS messages your bank sends you after a transaction and turns them into an actual transaction list. There's no open banking API here, so those messages are the only data source we have.

The thing doing the reading is 2,491 YAML files across 21 providers, compiled down to 78,410 lines of TypeScript.

This post is what's in them, and mostly why. A good chunk of the "why" is just me getting it wrong the first time and having to go back.

Quick note before we start: the pattern file below is the real one, and the message it reads is a real transfer of mine, twenty birr to a friend. As a rule the corpus doesn't leave the server, and a guard fails the deploy build when corpus text shows up in anything published. This file is the one exception I've carved out of it, because the made-up example I tried first kept flattening exactly the details worth showing.

What I'm actually parsing

21 banks and wallets, all of them a bit special, and all of them full of tiny variations that kept breaking my strict matching.

Telebirr alone has 501 pattern files. Not because telebirr does 501 different things, but because "you sent money" and "you sent money and paid a fee" and the Amharic version of each are separate strings with the fields sitting in different places.

Then there's everything in an inbox that isn't a transaction. OTP codes, promos, balance reminders. I have to be as good at throwing those away as I am at reading the real ones, because a promo misread as a transaction puts a fake number in someone's financial history.

Why these are YAML files

They used to be TypeScript. Moving them to YAML is the change that made this maintainable, for three reasons.

A pattern is data, and data wants a diff you can read. When a bank changes a template and I have to change a regex, the review should be the regex before and the regex after, on two lines. YAML is just funky formatted text, so that's exactly what the diff is: nothing else lives in the file to drag into it.

Tools that aren't the compiler can read data. Because a pattern is a file with a known shape, I've got a local admin app that loads it, runs a candidate regex against the real message corpus, shows me what starts matching and what stops, and writes the file back. That loop is most of why the corpus is any good, and I don't get it when the definition is code.

Generated code can't be edited in place. YAML is the source of truth, src/generated/ is output. If you hand-edit a generated file, the next codegen run quietly eats it. Making that whole tree obviously machine-owned takes away the temptation.

One pattern from start to finish

Say I send a friend twenty birr on telebirr. A few seconds later this lands. On a big enough screen it stays pinned while you read, and the parts being read light up.

Dear Robel
You have transferred ETB 20.00 to dagim alemu (2519****5183) on 12/09/2026 00:08:40. Your transaction number is DIC6NJHLOC. The service fee is  ETB 0.87 and  15% VAT on the service fee is ETB 0.13. Your current E-Money Account  balance is ETB 6,127.70. To download your payment information please click this link: https://transactioninfo.ethiotelecom.et/receipt/DIC6NJHLOC.

Thank you for using telebirr
Ethio telecom
Reading now: the identifier's whole claim

TRANSFER_TB_FROM_TB, the file that reads it, top to bottom.

name: TRANSFER_TB_FROM_TB # what every tool calls this pattern
provider: telebirr
category: transfer # money categories get the full health treatment; promos don't
order: 184 # tiebreak for when two patterns match the same body
status: done # triage state: draft, pending, flagged, done or junk
identifier:
  pattern: '(?=[\s\S]*VAT)(?=[\s\S]*service fee is\s+ETB)You have transferred ETB [\d,.]+\s+(?!Tip\b)to [\s\S]+?Your transaction number is'
  flags: is

identifier is separate from the field regexes. It decides whether this pattern applies at all, and everything below only runs once it's said yes. The two lookaheads at the front are doing real work: telebirr sends a nearly identical template where fee and VAT are one combined line, and that one is a different pattern. (?!Tip\b) keeps tips out of this one too.

regexes:
  SENDER_NAME:
    # The first capture group is the value that gets kept.
    pattern: '^Dear\s*([\w\s]+?)\s*You '
    flags: i
    source: common/NAME_PATTERN
  AMOUNT:
    pattern: transferred ETB\s*([\d,]+(?:\.\d+)?)
    flags: i
    source: common/TRANSFERRED_ETB_PATTERN
  TXN_ID:
    pattern: transaction number is\s*([A-Z0-9]+)
    flags: i
    source: common/TBTXN_PATTERN

Fields are named, not positional. Everything downstream says TXN_ID, so adding a field later doesn't renumber anything. And source means the regex is a reference to a shared copy: 64 of them are common enough to live in one file every pattern can point into.

  RECIPIENT_NAME:
    # Grabs everything up to the phone number or the date. The slot holds
    # whatever the other side's account says, junk included, so a character
    # class can't be written for it.
    pattern: \bto\s+([^(]+?)(?:\s*[\(]|\s+(?:on|at)\s)
    flags: i
  RECIPIENT_PHONE:
    pattern: (?:\(\s*(\+?[\d*]+)\s*\)|to\s+(\d{10,})\s)
    flags: i
  DATETIME:
    pattern: (?:on|at)\s*([\d/:-]+\s[\d:]+)
    flags: i

That loose capture on RECIPIENT_NAME is safe, because a name is never used to identify an account downstream. Only a number is, and the phone lands in its own field, mask and all.

  SERVICE_FEE:
    pattern: (?:The\s+)?service (?:fee|charge)(?:.*?)?is\s*ETB\s*([\d,]+(?:\.\d+)?)
    flags: i
  VAT_AMOUNT:
    pattern: (?:\d+%\s+)?(?:VAT(?:\s+is\s+ETB|\s+on\s+the\s+service\s+fee\s+is\s+ETB))\s*([\d,]+\.?\d*)
    flags: i
  BALANCE:
    pattern: balance is ETB *([0-9,]+(?:\.[0-9]+)?)
    flags: i
  RECEIPT_URL:
    pattern: (https:\/\/transactioninfo.ethiotelecom.et\/receipt\/[^\s.]+)
    flags: i
    fallback:
      template: https://transactioninfo.ethiotelecom.et/receipt/{TXN_ID}

Fee and VAT capture separately and get summed downstream: this message cost me 20 birr plus exactly 1.00 in charges, 0.87 fee and 0.13 tax. The fallback on RECEIPT_URL rebuilds the link from the transaction number when a message doesn't carry one.

extraction:
  direction: OUTGOING # from the account owner's side: money left
  amounts:
    principal: AMOUNT # the money that moved
    fees:
      - SERVICE_FEE
      - VAT_AMOUNT
    balance: BALANCE # what was left after
  participants:
    - fields:
        name: SENDER_NAME
      role: SENDER
      type: CUSTOMER
    - fields:
        name: RECIPIENT_NAME
        phone: RECIPIENT_PHONE
      role: RECEIVER
      type: CUSTOMER
  transactionIds:
    telebirr: TXN_ID
  timestampField: DATETIME

extraction is where captured text stops being strings and becomes a transaction. Participants carry roles: SENDER is me, RECEIVER is dagim. That's what turns a list of phone numbers into a list of people.

The schema catches my typos

Every file gets validated on the way in:

export const PatternDefSchema = z.object({
	name: z.string(),
	provider: z.string(),
	category: z.string(),
	order: z.number().default(0),
	status: z.enum(['draft', 'pending', 'flagged', 'done', 'junk']).default('draft'),
	identifier: RegexDefSchema,
	regexes: z.record(z.string(), RegexDefSchema).default({}),
	extraction: ExtractionSchema,
});

Misspell a role, point principal at a field that doesn't exist, name a debt product that isn't in the registry, and the build stops. Otherwise you get a pattern that runs fine and extracts nothing, which is a much worse afternoon.

If you make the wrong thing impossible to write down, you stop having to catch it later.

Before any regex runs

A message goes through three layers, in this order.

Layer one is an ignored set. 95 exact message bodies that get dropped without a word. Single characters, a bare account number with no sentence around it, truncated sends. It's a set lookup on the whole body so it costs nothing.

Each of these could have been a pattern instead. But there's nothing in them to extract, and they barely vary: a resend lands a character or two away from the last one. A pattern here is one more file in the 2,491 and one more identifier to keep from colliding, spent on a message I never wanted to parse in the first place.

The set is silent though. A body that goes in disappears without a trace, so I only add one I'd be fine never showing a user again.

Layer two is an exact-body map. Promos go out in bulk and are byte-identical to everyone. There's nothing in them to extract and no reason to run a regex, so they sit in a Map keyed on the full body and a hit returns immediately.

This one exists because the thing it replaced was actively hurting me. Promos used to get caught by broad catch-all patterns, one per provider per language. Those matched on the shared sign-off text at the bottom, usually the bank's own name, which every transactional template also has. My collision counter hit 373,000 messages matching more than one pattern. That's about 78% of everything I supported at the time.

Deleting the catch-alls and matching those bodies exactly fixed it in one go. Nobody writes these entries by hand. A harvester Claude wrote walks the corpus instead: group by body, keep anything with 2+ duplicates, skip what a specific pattern already handles, write out YAML.

Layer three is the regex patterns. Whatever's left is templated transactional text with fields that move, and that's what the 2,491 files are for.

Identifiers have to stand on their own

The rule that keeps this from collapsing: an identifier has to be right on its own, without depending on which pattern got tried first.

Leaning on ordering works the day you write it and breaks the day someone inserts a pattern above yours. And it breaks silently. The message gets filed as the wrong type, the fields still extract because the shapes are close enough, and nothing anywhere errors.

So identifiers use unique phrases plus negative lookaheads, the (?!Tip\b) from earlier being one. A transfer pattern says the thing only a transfer says, and explicitly says it isn't the tip or the reversal message sharing most of the same words.

There is still an order field, but it's a tiebreak, not a strategy. It's also where a genuinely annoying bug lived. 143 patterns share order: 150, and my sort had no second key, so when two tied the winner came down to whatever order the filesystem handed the files back in, which could differ between my laptop and the server.

The fix is one clause:

patterns.sort((a, b) => a.order - b.order || a.name.localeCompare(b.name));

Two patterns matching the same body is still a bug, though. It gets counted, not resolved by precedence, and the health sweep raises it as COLLIDES.

Faking a field that isn't there

Some fields aren't in the text but can be rebuilt from the ones that are.

You saw the real case above: telebirr's receipt URLs are just a fixed host with the transaction number stuck on the end. Messages from before a 2025 format change don't carry the URL at all, but they do carry the transaction number, so the link can be rebuilt exactly.

fallback:
  template: https://transactioninfo.ethiotelecom.et/receipt/{TXN_ID}

It only fires when the regex missed and every field named in the template matched. That second condition is the whole safety property. If a placeholder doesn't resolve you get nothing, instead of a URL with the literal text {TXN_ID} sitting in the middle of it pointing at a page that doesn't exist.

This template is also where a second product came from.

The page behind that URL is the provider's own statement of the transaction. My parse of the SMS can be wrong; the provider's receipt page can't. And if the URL is just a host plus a transaction number, you don't need the SMS at all to get there.

That's useful to people who've never heard of Odit. A transfer screenshot is easy to fake and the page the bank serves is not, so when someone claims they've paid you, the receipt link settles it. It started at v.odit.et as a page you pasted a receipt link into; it outgrew the subdomain and it's links.et now, reading receipts from 18 providers.

Codegen, and hashing everything into a version

npx tsx src/codegen/index.ts reads every YAML file and writes src/generated/. One patterns module and one exact-body module per provider, the shared regexes, the ignored set, and a version.

The version is the fun part:

const hash = crypto.createHash('sha256');
// _common/regexes.yaml, then _ignored.yaml, then every _exact/*.yaml,
// then every provider directory, each file list sorted.
return hash.digest('hex').slice(0, 16);

PATTERN_VERSION is a content hash over everything that can change a parse result. Nobody bumps it by hand, so nobody can forget to.

Every directory read gets sorted before it's hashed, because readdirSync order depends on the filesystem, and an unsorted hash would come out different on two machines holding identical files. A version that changes when the content didn't is worse than having no version at all, because it kicks off a full reparse of the corpus for nothing.

It also has to cover the ignored set and the exact-body maps, not just the regex patterns. I missed that at first. An earlier version of this hash only read the provider directories, so adding a promo body to the exact map produced no version change, nothing reparsed, and the messages it should have reclassified just sat there as unknown until something unrelated forced a run.

The page I actually live in

None of the above gets you a good corpus on its own. What does is that every pattern gets run against a few million real messages before I commit it.

The cheap version of that is just throwing a candidate regex at the message table, which I can do because definitions are data:

select raw_body from sms_export_items where raw_body ~ $1 limit 200

That tells you what a regex catches, and more usefully what it catches that you didn't mean. But it doesn't tell you whether the pattern's fields came out, and that's where the actual bugs are. A pattern can win on a message and still leave the balance empty, and you'd never know from a match count.

So there's a page per pattern in the local admin app, and the top of it is a whole-corpus extraction verdict. For TRANSFER_TB_FROM_TB it currently reads:

SWEPT   100.0% of 115,034 extract cleanly
        missing: VAT_AMOUNT 2   SENDER_NAME 2   BALANCE 1

Wins 115,034 messages from 664 people. Extracts fully on 99% of wins;
misses VAT_AMOUNT on 2, SENDER_NAME on 2, BALANCE on 1.
Derives RECEIPT_URL from TXN_ID on 2 wins.

Five bad messages out of 115,034. I'd never have found those by paging through dots.

The failure shapes

The bit that saves the most time is that it doesn't list the five failures. It lists the shapes:

FAILURE SHAPES · ONE REAL MESSAGE EACH
  2x  SENDER_NAME
  1x  VAT_AMOUNT BALANCE
  1x  VAT_AMOUNT

Three distinct ways this pattern can come up short, one real body shown for each. Fixing a shape fixes every message in it, and the count tells me whether it's worth touching the regex at all. Two messages out of 115,034 usually isn't.

Next to it sits the richest fully-clean body as a reference, with every captured span highlighted in the message itself, so you can see exactly what each regex grabbed rather than trusting that it did. On my transfer from earlier that's:

SENDER_NAME      Robel
AMOUNT           20.00
RECIPIENT_NAME   dagim alemu
RECIPIENT_PHONE  2519****5183
DATETIME         12/09/2026 00:08:40
TXN_ID           DIC6NJHLOC
SERVICE_FEE      0.87
VAT_AMOUNT       0.13
BALANCE          6,127.70
RECEIPT_URL      https://transactioninfo.ethiotelecom.et/receipt/DIC6NJHLOC

Seeing the spans is how you catch a regex that matched the wrong thing. A lazy capture that runs one word too far still counts as a match, and it still looks green on any pass-fail check.

What counts as a miss

The verdict rule is in field-verdict.ts:

a message is green when every REQUIRED field either captures or renders from its fallback template, red when none do, yellow in between. Optional fields never count against a body; uncompilable regexes are invisible.

Required means !optional && !missVerified. RECEIPT_URL stays required here because its fallback can always rebuild it. A field with no fallback that the sender genuinely omits gets optional instead, or every old-format message sits here as a permanent miss and the three real shapes drown under thousands of fake ones.

missVerified is the same idea with a human behind it: I looked, the sender genuinely omits it, stop counting it against me.

And "Derives RECEIPT_URL from TXN_ID on 2 wins" is the fallback template doing its job in the wild. Those two messages don't carry a URL, the template rebuilt it from the transaction number, and they counted as clean rather than as misses.

Every field captures before anything is judged, because a required field's fallback template is allowed to reference an optional field's value. Judge as you go and the derive silently fails.

That same rule is used in three places, on purpose. The dots down the side of the page, the API's failing-first scan, and the health sweep's accumulator all call the same function, so they can't quietly disagree about what a broken message is. The page also shouts index predates current patterns when the sweep it's reading is older than the YAML, because a green bar built from last week's regexes is worse than no bar.

The whole app only runs on localhost and it's excluded from every deploy, because it's a tool that puts production regexes and real message bodies on screen.

The health sweep

This is the piece that keeps a corpus this size from rotting without me noticing.

It walks the whole corpus once and prices every detector in that single pass, then flags patterns:

FlagWhat it means
BAD-IDENT / BAD-REGEXthe regex doesn't compile
COLLIDEStwo identifiers match the same body
DEADzero matches in the entire corpus
SHADOWEDit matches, but another pattern usually wins
UNPARSEDan amount field grabbed something non-numeric
P-LOW / B-LOWprincipal or balance under 98% coverage
PARTIALover 10% of wins only extract some required fields
RESIDUEmoney-shaped tokens left uncovered on a lot of wins
DIVERSEone pattern matching loads of different body shapes
AUTO-ROLEa participant role never got triaged

DEAD and DIVERSE are the two that keep earning their keep.

DEAD catches a pattern I wrote for a format the bank has since dropped. It compiles fine and it looks maintained, so nothing about the file tells you it stopped matching anything months ago. You only find out by counting across the whole corpus.

DIVERSE catches the opposite: an identifier loose enough to match free text, winning across a pile of differently-shaped bodies. That's a pattern turning into a catch-all, which is the exact thing the exact-body layer was built to undo. Nice to catch it before the collision count blows up again.

Two flags can be silenced per pattern, and both need me to say so in the file. residueVerified when I've looked at the leftover tokens and they're carrying nothing, and missVerified when a field is genuinely missing upstream rather than being missed by me. Measurement keeps going either way, only the flag stops.

The output gets compared against a committed baseline, so making things worse shows up as a diff instead of as a support message six weeks later. Accepting a new baseline is a commit, which makes it a decision somebody made instead of drift.

What's left

  • Content-hash identity end to end, so the row id stops being load-bearing
  • Auto-proposing YAML for unmatched clusters instead of me writing each one
  • A proper story for a provider changing a template mid-flight, rather than waiting for DEAD to tell me weeks later
  • The Amharic patterns are thinner than the English ones and I know it

The user side of all this is that they don't do anything. Install the app, grant message access once, and the history builds itself out of what's already on the phone, including everything from before they installed it. They never download a statement and they never hand a password to anybody.