← all work

research tool · 2026

Quran Malay Audit

A small audit toolkit that started with a few typos I found while building goSolat's Quran module.

view on github

I wasn’t trying to audit a Quran translation.

I was building the Quran module for goSolat.

While preparing the Malay translation for the app, I started noticing things that didn’t look right. A few spelling errors. Inconsistencies. Small things that were easy to dismiss, until I started checking them against other sources.

Then I realised I couldn’t just fix them.

I needed to know which source I was looking at, which edition it came from, when it was retrieved, and whether the difference was actually an error or simply another translator’s choice.

So I started digging.

And the deeper I went, the more uncomfortable the question became:

How long had some of these errors been sitting there?

The Malay translation most commonly encountered online comes from Tafsir Pimpinan al-Rahman, associated with Sheikh Abdullah Basmeih. It is an important work, but the digital data I was working with was not simply a matter of copying an authoritative book into an app. Different providers carried different versions, metadata was sometimes inconsistent, and some text needed closer review.

What surprised me most was not that errors existed.

It was how easy they were to leave untouched.

People with deeper knowledge of Quranic Arabic, tafsir, and translation can often recognise problems, consult the original Arabic, compare tafsir, or simply use another translation. The ordinary reader usually doesn’t have that option.

So a small typo I found while building an app turned into a much bigger question:

Who checks the translation that everyone else is consuming?

That became Quran Malay Audit.

Not a translation rewriting tool

I didn’t want to build another system that simply compares two texts and declares one of them wrong.

A difference between two translations is not automatically an error.

Every finding therefore stays attached to its source key, provider, resource ID, translator, retrieval date, version, and evidence. Quran.com resource 39 and QUL resource 292 remain separate tracks, even when they contain similar Malay text.

The source is part of the data.

First, make sure the data is actually complete

The toolkit accepts key-value JSON, nested arrays, record lists, translation lists, and SQLite exports.

It checks them against the canonical 6,236 Quran verse keys and reports:

  • missing keys
  • duplicate keys
  • unexpected keys
  • empty translation text

It doesn’t rewrite the source.

A complete export only tells me that the data is structurally complete. It doesn’t tell me that the translation is linguistically correct.

That distinction became important very quickly.

Compare without declaring a winner

The comparison command produces source-scoped candidate findings with the verse key, observed text, proposed text, source identities, and evidence.

A difference can be:

  • a spelling issue
  • punctuation
  • formatting
  • a genuine translation difference
  • or something that needs someone with the right expertise to look at it

So a comparison finding is explicitly marked as a candidate, not an error.

Two translators can make different choices. The tool records the evidence and leaves the editorial decision to a reviewer.

The workflow is deliberately boring:

validate → compare → review → confirm → apply

That’s the point.

Corrections have a gate

Only a confirmed correction from the matching source track can be applied.

The application step checks the exact verse key, source key, expected text, whole-word boundaries, and optional occurrence count. If the text has already been corrected, the operation is idempotent. If the precondition is missing or ambiguous, it stops.

I wanted the software to make it difficult for me to accidentally “correct” something that I had no business correcting.

A bounded advisory AI pilot

The latest iteration adds a small standard-library client for Jev advisory review.

It sends one candidate finding at a time with only the fields needed for review: source identity, resource ID, verse key, observed text, proposed text, and evidence.

Jev can suggest the likely kind of issue and whether specialist review may be needed.

But it cannot approve anything.

The result is written to a separate report marked advisory_only: true. It cannot promote a finding, modify a manifest, or rewrite source text.

AI helps with the queue.

It doesn’t become the authority.

What started as a typo became a provenance problem

This project started because I wanted clean Quran data inside an app.

It ended up making me think much more carefully about what happens when religious text becomes digital infrastructure.

Once a translation is copied into APIs, databases, apps and websites, an old mistake can become very easy to reproduce and very hard to notice.

And there is an interesting gap here.

People who study the Quran deeply have ways of checking, comparing and contextualising translations. Most people using a Quran app simply see the text presented to them.

They trust that someone has already checked it.

Sometimes that assumption is reasonable.

Sometimes it isn’t.

I don’t think software should pretend to solve that problem by itself.

What software can do is make the trail visible:

where did this text come from, what changed, what evidence do we have, and who actually approved the change?

That is what Quran Malay Audit is for.

And that is probably the most important thing I learned while building the Quran module for goSolat.

Python · Data validation · Source provenance · Human review · Open source

Python · Data Validation · Open Source · Research · Content