How Our Tajweed Check Actually Works — And What It Can't Do
Most apps that claim to check your tajweed do not tell you how. This is an account of exactly how ours works, including the parts where it falls short — because a tool that judges recitation of the Qur'an should be auditable, and because you cannot sensibly trust feedback whose limits you do not know.
The short version: we measure three rules that have clear acoustic signatures, we report how confident we are, and we say plainly that the rest needs a teacher. We call it a best-effort tajweed rating, and the estimate we publish is that it gets you roughly 60% of the way. This article is the long version.
Why most of tajweed is not measurable
Tajweed is a discipline of precision. It governs the articulation point of each letter (makhraj), the qualities that letter carries (sifat), how long vowels are held, when sounds merge, when they are hidden, when the voice is nasalised, and where a reciter may stop. It is transmitted through talaqqi: reciting to a qualified teacher who hears the error and corrects it, in a chain reaching back to the Prophet ﷺ.
Software hears none of that. Software hears amplitude over time. To check a rule automatically, the rule has to leave a fingerprint in the waveform that a machine can find reliably — and most tajweed rules do not, at least not without acoustic models trained specifically on Qur'anic Arabic that do not exist in open form.
So instead of building something that gestures at all of tajweed and gets most of it wrong, we picked the rules that are genuinely measurable and left the rest alone.
The three rules we check
Madd — elongation
A madd is a vowel held longer than a beat. A natural madd (madd tabee'i) runs about two harakat — two beats of a short vowel. Others run four or six.
Duration is the easiest thing in the world to measure, provided you know where the word starts and ends. When you recite into the app, the recording goes to Whisper for transcription, and we ask for word-level timestamps along with the text. That gives us a start and end time for each word.
Here is the important limitation, stated up front: those are word boundaries, not phoneme boundaries. We can see that a word containing a madd took less time than it should have. We cannot see which syllable inside it was clipped.
That makes this a detector for the gross error — not elongating at all, or stretching a two-beat madd to six — which happens to be exactly the mistake beginners make. It is not a detector for a madd that is slightly short.
We also do not compare against any fixed words-per-minute constant. The app works out your own pace from the median across every word in that recording, then measures the madd words against your baseline. A slow reciter and a fast one are judged the same way. If a recording is too short to establish a baseline — fewer than three usable words — we report that we could not measure it rather than guessing.
Ghunnah — nasalisation
Ghunnah is the nasal sound held for about two beats on a noon or meem carrying shadda, and in idgham and ikhfa.
Nasalisation has a real acoustic signature. When the nasal cavity is coupled into the vocal tract, energy concentrates in a low nasal formant around 250–350 Hz, and anti-formants damp the region between roughly 1 and 2.5 kHz. So the ratio of low-band to mid-band energy jumps, and stays raised for as long as the sound is held.
The app computes that ratio across short overlapping frames of your recording and looks for stretches where it stays elevated for at least 100 milliseconds — long enough to be a held ghunnah rather than a nasal consonant in passing.
Qalqalah — the bounce
Qalqalah is the echoing release on ق، ط، ب، ج، د when they carry sukun. Acoustically it is a plosive: a closure where energy drops close to silence, followed by an abrupt burst as the articulators release.
That shape — a quiet minimum followed within a few tens of milliseconds by a steep rise — is straightforward to find in an energy envelope. The app looks for closures that are genuinely quiet relative to that recording's own median level, so a quietly recorded session is not read as one long closure.
Where the rules should apply
Measuring sound is only half of it. The app also has to know where each rule ought to occur, and that half is pure text analysis over the Uthmani script — completely deterministic, no model involved.
For any ayah, the app walks the letters and their diacritics to find:
- Madd sites — an alif after a fatha, a waw after a damma, a ya after a kasra, each with no vowel of its own; plus the superscript alif, which is the same sound with no written letter; plus anything carrying the maddah sign, which marks the longer obligatory madds.
- Ghunnah sites — noon or meem with shadda (always ghunnah); a saakin noon followed by ي ن م و (idgham with ghunnah); a saakin noon followed by one of the fifteen ikhfa letters. Note that the letter triggering the rule is often in the next word, so this analysis runs across the whole ayah rather than word by word.
- Qalqalah sites — ق ط ب ج د carrying sukun.
One deliberate omission: a qalqalah letter at the end of an ayah also bounces when you stop there — but whether you stop is your choice, and marking it as required would penalise a correct continuous recitation.
This text analysis runs against the full committed Qur'an corpus in our test suite, so the rule detection is exercised against all 6,236 ayahs as they actually appear, not against simplified examples.
Combining it into a rating
The score is a plain percentage of the assessed rules that passed. Not a weighted blend — a weighted blend would imply a precision these detectors do not have, and would let a confident qalqalah result paper over a madd that was never measured.
The single most important design rule in the whole system is this: bias toward silence.
If an app tells someone their recitation is correct when it is not, that error gets reinforced at every spaced-repetition interval for years. That is a different category of harm from a language app mishearing an accent. So:
- A rule that could not be measured reports "not assessed" — never a passing mark.
- Missing word timings, or audio the browser could not decode, produce no score at all rather than a zero.
- The wording throughout describes what was heard, never what is right.
Confidence never reads "high". Even when all three rules are measured, letter-level articulation is still completely unexamined — and the rating says so.
The part we did not build, and why
We planned three approaches. Two shipped. The third did not, and it is worth explaining because it is the one that would have made the biggest difference.
Forced phoneme alignment means aligning the expected phonemes of the ayah to the audio and grading each one — this is how you would actually check the makhraj of a letter, or whether ص was pronounced distinctly from س. It is what serious tajweed research works on.
It requires a CTC acoustic model fine-tuned on Qur'anic Arabic phonemes, and a labelled corpus to fine-tune it on. Neither is something we can improvise.
We could have shipped something that looked like it: align Whisper's output characters to the expected text and emit per-letter verdicts. It would have produced confident, specific, letter-by-letter feedback with no acoustic basis whatsoever. That is worse than not having the feature, because it would be believed.
So the slot is empty, the confidence never reaches "high", and the rating states that individual letters are not checked at all.
What this means for you
Use the rating for what it is: a fast check on three things that are easy to get wrong and easy to measure. If it says your elongations are being cut short, they probably are, and that is genuinely useful to know between lessons.
Do not use it as a verdict on your recitation. It has not examined your letters. It cannot hear the difference between a heavy and a light ر. It does not know if your ض is coming from the right place. These are not oversights — they are the parts of tajweed that require a human being who has been taught by a human being.
The right way to use a tool like this is as practice between sessions with a teacher, not as a replacement for one. Recite to a qualified reciter. Get corrected. Then use the app to keep those corrections honest during the week.
Why we published the method
Two reasons.
The first is that people are entitled to know how something judging their recitation of the Qur'an works. Feedback you cannot audit is feedback you cannot calibrate, and calibration matters enormously here — knowing that we do not check letters tells you exactly how much weight to give a high score.
The second is that a claim about tajweed should be falsifiable. Everything described above is deterministic and testable. The signal detectors are verified against synthetic tones, the rule detection against the whole Qur'an. If we have made an error, it is an error someone can find and point at, which is the only kind worth making in this domain.
If you want to check our work, the method is exactly as described here: three rules, two measurement techniques, one honest gap where the hardest problem sits unsolved.
Frequently asked questions
Can an app actually check your tajweed?
Partly, and only for some rules. Certain tajweed rules have a clear acoustic signature that software can measure — how long an elongation is held, whether a nasal sound is present, whether a letter is released with a bounce. Others, especially the precise articulation point of individual letters, need either a trained phoneme model or a human ear. Any app claiming to check tajweed completely is overstating what it does.
What tajweed rules does The Hifz Project check?
Three: madd (elongation), ghunnah (nasalisation), and qalqalah (the bounce on ق ط ب ج د at sukun). These were chosen specifically because each one maps onto something measurable in a recording. Rules about heaviness and lightness, the makhraj of individual letters, and the finer points of stopping are not checked at all.
Does the app replace a teacher for tajweed?
No, and it is designed not to pretend otherwise. The rating is labelled best-effort, reports roughly 60% coverage, and names what it did not measure every single time it appears. Tajweed is transmitted through talaqqi — reciting to a qualified teacher who corrects you — and no scoring function substitutes for that chain.
Why does the app say 'not assessed' instead of giving me a score?
Because a missing measurement is not a passing grade. If the recording could not be decoded, or the transcription came back without word timings, the honest answer is that nothing was measured. Reporting that as a score of zero would be wrong, and reporting it as a pass would be worse — a learner told their recitation is correct when it isn't will reinforce the error at every review for years.