---
title: "Why AI gets Arabic calligraphy wrong"
description: "One model reads Arabic handwriting at under one per cent word error. The best anyone has managed on classical calligraphy is forty-five. I went looking for the missing data that would explain the gap and found the opposite — a script whose rules were written down in the tenth century, and again in software in the 1980s."
author: "Omar Albeik"
date: 2026-08-07
type: article
topics: [calligraphy, typography, ai, software, engineering]
language: en
reading_time_minutes: 29
canonical_url: https://omaralbeik.com/en/blog/why-ai-gets-arabic-calligraphy-wrong
translation_url: https://omaralbeik.com/ar/blog/why-ai-gets-arabic-calligraphy-wrong
source_url: https://omaralbeik.com/en/blog/why-ai-gets-arabic-calligraphy-wrong.md
---

# Why AI gets Arabic calligraphy wrong

Type a line of Arabic into any of the calligraphy generators and you get back
something with the correct temperature and no content. The strokes have weight.
The composition stacks the way Thuluth stacks. Somebody who has never read a
word of Arabic would accept it without a second look, and that is exactly the
population it was built for.

Anyone who reads Arabic sees it in under a second. Letters that should have
joined are sitting apart. A ه has taken its initial form in the middle of a
word. The dots have drifted off the letters they belong to and settled wherever
the density looked right. It is not bad calligraphy. Bad calligraphy is a thing
a beginner produces, and it is legible. This is calligraphy-shaped noise.

I wanted to know why. I build software for a living and I have spent enough
years failing to write a decent *alif* to know roughly what the craft asks of
you, which is a strange pair of interests to have but a useful one for this
particular question. What follows is what I found looking into it, including
the part where the explanation I assumed was true turned out to be wrong.

<PostFigure
  src={thuluthYaqut}
  alt="A calligraphic fragment in Thuluth script by the 13th-century calligrapher Yaqut al-Musta'simi, with dense interlocking letterforms and dots distributed through the composition."
  caption="Thuluth by Yaqut al-Musta'simi, the thirteenth-century master whose hand is still a reference seven centuries on. Every relationship in this panel — how far a stroke travels before it turns, where a dot sits relative to the letter that owns it, which words climb and which stay on the line — is a decision, and the ones the generators get wrong are all here. Library of Congress, public domain."
  locale="en" eager
/>

## Not for want of data

My first assumption was the obvious one: Arabic is under-resourced, and the
models will catch up. That was true for a long time. It has stopped being true,
and the numbers now say something stranger.

[Qalam](https://arxiv.org/abs/2407.13559), a vision encoder-decoder trained on
four and a half million manuscript images, reads Arabic handwriting at **0.80%
word error**. [QARI-OCR](https://arxiv.org/abs/2506.02295), a 2B-parameter
Qwen2-VL fine-tune, reads diacritically-rich Arabic print at **6.1% character
error** — with the tashkeel, which is the hard part. Both are open. Both run on
hardware you own.

Three numbers recur below, so it is worth fixing them now. Word error rate is
the share of words a system gets wrong; character error rate the share of
characters, which is the gentler measure since a word can be wrong by one
letter. Exact match is the strict one: the share of images transcribed with
nothing wrong at all.

Then somebody built a benchmark for calligraphy specifically.
[DuwatBench](https://arxiv.org/abs/2601.19898) landed in January, out of
MBZUAI: 1,272 samples across Thuluth, Diwani, Kufic, Naskh, Ruq'ah and
Nasta'liq, with transcriptions, style labels and bounding boxes. Thirteen
models were pointed at it, open and closed, general-purpose and Arabic-specific.

<PostFigure
  src={duwatStyles}
  alt="A table showing the basmala written in Thuluth, Diwani, Kufic, Naskh, Ruq'ah and Nasta'liq, alongside sample counts of 706, 230, 83, 110, 76 and 67 respectively."
  caption="One phrase — بسم الله الرحمن الرحيم — in each of the six styles the benchmark covers, with how many samples of each it holds. These are the same four words every time. A model that has learned Naskh has not learned Thuluth, and the sample counts show which end of that range the data sits at. From DuwatBench (Patle et al., 2026), CC BY 4.0."
  locale="en"
/>

The best result on the board is Gemini 2.5 Flash, at **45% word error and 42%
exact match**. Across the thirteen models tested, exact match runs from **zero**
to that 42%, and averages **17.9%**. The
open models are worse in ways that are almost funny to read about: one of them
responds to difficulty by emitting the word *Allah* over and over, having
correctly learned that this is a good bet on Arabic calligraphy and nothing else
about the image in front of it.

<PostFigure
  src={duwatPredictions}
  alt="A large comparison grid: eleven calligraphic images with ground-truth transcriptions, followed by the predictions of thirteen open- and closed-source models, with correct answers highlighted in green and most cells incorrect."
  caption="Eleven images across the top, the correct reading beneath them, and then every model's attempt. Green is a hit. Read down the InternVL row and watch it answer الله to almost everything; read the TrOCR row and watch it return a single dot. GPT-4o's contribution to several columns is a refusal — it declines to help, in Arabic. The models are not slightly wrong here. They are not reading. From DuwatBench (Patle et al., 2026), CC BY 4.0."
  locale="en"
/>

So a system built for Arabic manuscripts reads them at under one per cent word
error, and the best system anyone has pointed at a calligraphic panel gets 45%.
Nor is this general-purpose models being asked something outside their brief.
Two of the thirteen were Arabic-specific, and both sit low on the board:
Microsoft's Arabic handwriting TrOCR returns **99.98% word error and not one
exact match** across the entire benchmark. A model built for Arabic handwriting,
handed Arabic handwriting, reading essentially nothing.

I expected the per-style breakdown to be a clean ramp: plain scripts easy,
ornate scripts hard. It is not, and the way it fails to be is the most
interesting thing in the paper.

<Figure caption="Word error rate for each style, best closed model against best open one. Not a ramp. Kufic is worst for both, Naskh is the steadiest, and Thuluth — the ornate one — is where the closed model does best of all six, twenty-eight points clear of the open model on the same images. Numbers from DuwatBench Table 3 (Patle et al., 2026), CC BY 4.0.">
  <StyleErrorRates locale="en" />
</Figure>

Kufic being the hardest is not obvious. It is the most geometric of the six, the
one closest to something you could describe with a grid, and it defeats
everything, presumably because that same geometry strips out the redundancy a
reader normally leans on.

The Thuluth result I do not think is a reading result at all. Thuluth is 55% of
this benchmark and the corpus is 44% Qur'anic, so a model with a strong language
prior has a great deal of it effectively memorised, and guessing the basmala is
not the same skill as reading it. This looks at first like a contradiction with
the paper's own summary, which says Naskh and Ruq'ah are the most accurately
recognised styles. It is not. That holds across all thirteen models averaged
together, and the same passage notes separately that Thuluth and Diwani run
substantially higher for the closed-source models in particular. The Gemini
column is the second half of that sentence.

## Not for want of scale

The comfortable next move is to say this is still a data problem, just a bigger
one. A paper from June puts that to rest.

Sana Al-azzawi, Elisa Barney and Marcus Liwicki ran [one model across nine
handwriting datasets](https://arxiv.org/abs/2606.18884) — Arabic, Urdu, Persian,
English, German — holding the architecture fixed and varying only the script and
the volume of training data. The Arabic-script penalty is largest when data is
scarce, narrows as data grows, and then **stops narrowing at a consistent five to
seven character-error points**. Cleaning the annotations helps and does not close
it.

Their diagnosis is the thing you notice within about ten minutes of picking up a pen. **Around 30% of the substitution
errors on Arabic-script data are confusions between visually similar characters,
against about 15% on Latin.** One group of five letters in this alphabet is told
apart by nothing
but dots. Standing alone they differ a little; inside a word they collapse onto
the same stroke and the dots are the whole of what is left. In a Latin alphabet
the letterform carries the identity and the diacritic modifies it. Here the
diacritic *is* the identity, riding on a stroke that means nothing by itself.

<Figure caption="Five letters in the middle of a word. There is no character for the top row — it had to be drawn, because no font will give you a letter with its dots taken off. That is the point: the shared stroke is not a letter, it is the part five letters have in common, and everything that makes it a particular word is the marks above and below it.">
  <DotIdentity locale="en" />
</Figure>

Now put that in a Thuluth composition, where the calligrapher moves dots off
their letters to balance the block — a licensed, traditional, expected move that
every trained reader parses without effort because they are reading the *word*,
not the glyph. The model is reading glyphs. It was always reading glyphs.

<PostFigure
  src={kuficQuran}
  alt="A parchment folio from an eighth or ninth century Qur'an in early Kufic script, with wide horizontally-extended letterforms in brown ink and sparse markings."
  caption="A Qur'an folio in early Kufic, eighth or ninth century. Look closely at what the dots are doing: the letters carry no pointing at all — no i'jam, nothing to tell a bāʾ from a tāʾ — while the red dots scattered across the page are a separate, older system marking vowels. Arabic acquired its dots in layers, for different jobs, centuries apart. Public domain."
  locale="en"
/>

If you want the number that makes this unarguable, look at Ajami, the
Arabic-script orthographies of West Africa. The [first HTR dataset for Fulfulde
and Hausa manuscripts](https://link.springer.com/chapter/10.1007/978-3-032-04627-7_36)
was published in the ICDAR 2025 proceedings. Its authors ran a suite of
Arabic-script recognition models *built for historical manuscripts* over it, and
report **character error rates of 65 to 84%**. Not word error — character error,
the gentler of the two measures. Up to five letters in six wrong, one letter at a
time, by models built for exactly this kind of material. "Arabic AI" turns out to
mean a narrow slice of Arabic.

<PostFigure
  src={ajamiFodio}
  alt="A weathered manuscript page in a West African Arabic hand, four lines of large brown script with dense smaller annotations written between and above them, and a geometric decorated panel filling the lower half."
  caption="A page from a poem by Abdullahi dan Fodio, of the Sokoto Caliphate. Everything here is Arabic script and almost none of it is what a model trained on Cairo has seen: a West African hand, a second smaller hand crowding the glosses, its own orthographic conventions, a decorated panel that is not text at all. Before a model can misread this it has to work out which marks are even the line. British Library / Endangered Archives Programme, CC0."
  locale="en"
/>

## What the script is doing

At this point I stopped reading benchmark papers and went back to the script
itself, because the errors had started to look structural rather than
statistical.

Arabic is not a set of glyphs you substitute one for one. A letter takes a
different form depending on whether it begins a word, sits inside it, ends it, or
stands alone. Six letters — ا د ذ ر ز و — refuse to join to whatever follows
them, so a single word breaks into several connected runs. Certain pairs fuse
into ligatures that are not either letter. In the classical styles the line
itself stops being horizontal: Diwani sweeps, Thuluth stacks words vertically
into a block, and where a word sits is a compositional decision made by a person
weighing the whole panel.

<Figure caption="The two rules that a character-for-character swap cannot honour. Above: one letter, four shapes, chosen by its neighbours. Below: a word that is not one connected run but three, because dāl and rāʾ refuse to join to what follows them. Those gaps are not spaces, whatever a renderer might assume. Get either wrong and a native reader sees it immediately.">
  <JoiningRule locale="en" />
</Figure>

None of that is decoration on top of the letters. It *is* the letters. The script
has a grammar, and the grammar is the art.

<PostFigure
  src={thuluthAlphabet}
  alt="A calligraphic album folio showing the Arabic alphabet written out in Thuluth script, letters grouped and repeated across ruled lines."
  caption="The alphabet set out in Thuluth by 'Ala' al-Din Tabrizi. This is the exercise every student copies, and it is also, read the right way, a specification: the same letter appears in its different positions, the proportions are fixed, and what looks like a display of virtuosity is a table of rules. Dated 1593; Library of the Golestan Palace, via Reed College. Public domain."
  locale="en"
/>

And the line does not have to be a line. Diwani, the Ottoman chancery hand, packs its words into a rising sweep and closes the gaps until the block reads as one mass, a habit traditionally explained as a security measure, since a text with no room between its words is a text nobody can insert a clause into.

<PostFigure
  src={diwaniFirman}
  alt="A tall Ottoman imperial decree written in Diwani script, its lines rising diagonally in dense packed clusters with gold illumination at the head."
  caption="An imperial decree of Mehmed II in Diwani. The words climb, the spacing is deliberately eliminated, and the whole block is engineered to be unforgeable and unamendable. Every property that made it good chancery practice makes it hostile to a model: no baseline, no word gaps, no isolated forms to latch onto. Google Art Project, public domain."
  locale="en"
/>

Nasta'liq goes the other way. The name contracts *naskh* and *ta'liq*, and ta'liq means suspension, which is what the script does: words hang off one another and descend as the line runs, so a reader tracks a diagonal rather than a baseline. A layout engine that assumes text sits on a horizontal has already lost.

<PostFigure
  src={nastaliqPanel}
  alt="A Persian calligraphic panel in Nasta'liq script, with words descending diagonally across the page in flowing, hanging lines."
  caption="A Nasta'liq panel. There is no baseline here in any sense a text layout engine would recognise — words hang off each other and the eye follows the slope down and to the left. This is the style DuwatBench holds the fewest samples of — sixty-seven. Public domain."
  locale="en"
/>

A model trained on images of finished calligraphy is being asked to recover that
grammar from its residue. It sees where the ink landed. It never sees the rule
that put it there, or the pen angle, or the order the strokes were made in, or
the fact that the calligrapher moved that dot deliberately. Given enough
examples it learns the texture of the residue, which is precisely what it then
produces.

## The rules were already written down

Here is where I expected to find the gap — some vagueness in the tradition, a
body of knowledge too intuitive to commit to paper — and found close to the
opposite.

The proportions are not taste passed down as vibes. They are a ratio system, and
it was formalised more than a thousand years ago as *al-khaṭṭ al-manṣūb*,
proportioned script, traditionally credited to the Abbasid vizier and
calligrapher Ibn Muqla, who died in 940. It has three units.

<Figure caption="The dot is not punctuation here: it is the ruler. Press the cut nib once and you get a rhombus; stack a fixed number of them and you have the height of an alif, five of them in Naskh and a different count for each style; take that height as a diameter and you have the circle every curve is checked against.">
  <ProportionSystem locale="en" />
</Figure>

That system is why a trained reader can look at a letter and say it is
*incorrect* rather than just unfamiliar: there is a number it failed to meet.
Ibn Muqla is credited with canonising the six classical pens on the same basis,
and Ibn al-Bawwab and Yaqut al-Musta'simi refined the system after him. A
thousand years before anyone wrote a font engine, this script had a written
specification.

And it did not stop at the letter. Whole page layouts were fixed forms too,
reproduced by generations of calligraphers to the same plan.

<PostFigure
  src={hilyeHafizOsman}
  alt="An Ottoman hilye panel: a rectangular illuminated composition with a band of Thuluth at the top, a central circular medallion of Naskh text ringed in gold, four small corner medallions, and a framed verse below."
  caption="A hilye by Hafiz Osman (1642–1698), who devised the layout that every later hilye follows. This is a template: the basmala across the top, the round belly at the centre carrying Ali ibn Abi Talib's description of the Prophet, the four corner medallions naming the first caliphs, the verse band beneath. All in fixed relation, and every hilye written since sits inside those constraints. Sadberk Hanım Museum, public domain."
  locale="en"
/>

That is the tenth century. The modern version is the part that actually
reframed this for me.

[DecoType](https://decotype.com/) was incorporated in 1985 by Peter Somers,
Mirjam Somers and Thomas Milo — Designers of Computer-aided Typography. What
they built was not a glyph set but a font that knew the script's rules and
applied them: their Ruq'ah was the first dynamic Arabic typeface of its kind,
and by the early 1990s it was shipping inside Microsoft Office. The DecoType
Naskh and Thuluth families went out to a generation of Office users under an "MS
Office Edition" label, on licences dated 1992 to 1998. Whether current builds
still carry them I could not establish.

Underneath sat ACE, the Arabic Calligraphic Engine, built first for Ruq'ah and
then extended across a broad analysis of Naskh, and later renamed the Advanced
Composition Engine as it grew to drive any Arabic typeface.

Milo's [own account of the model](https://www.academia.edu/2448322/A_model_for_Handling_the_Arabic_Script_2005_)
is explicit about the stakes: ACE approaches Arabic typography in analogy with
*pre-typographic text manufacture*. It encodes the orthographic rules by which
letterform combinations weave into ligature forms. It is a considered rejection
of movable type as a model for this script, a rejection made by people who went
and studied the manuscript tradition first. He took the Peter Karow Award for it
in 2009.

It reached working designers as Tasmeem, the Arabic plug-in for InDesign ME that
DecoType built with WinSoft International, who published it. As far as I can tell it remains the most structurally faithful Arabic
composition system anyone has shipped.

WinSoft [discontinued it in April 2021](https://winsoft-international.com/tasmeem/) —
the product, not DecoType, whose engine had always been the part underneath.
Last release for InDesign CC 2021. Existing clients supported to the end of their
subscriptions, no new sales, no further upgrades. DecoType's own site now reads
as an archive; the most recent thing dated on it is a 2017 lecture.

One of the people who worked on that plug-in is still at it, and what he sells
now is the neatest summary of where this has ended up. Nihad Taisir Nadam —
Damascus-born, Dubai-based — co-developed the first online Arabic digital
calligraphy tool with WinSoft in 2011, alongside Tasmeem itself, and by his own
account has worked with Adobe, WinSoft and DecoType on Arabic font technology.
He now sells [ArabicDesign.ai](https://arabicdesign.ai/en), which generates
calligraphy from typed Arabic in about fifteen styles and exports SVG.

Read its feature list closely. It offers to *preserve the requested text at up
to 95%*, and to *reduce* common letter and dot mistakes. That is a purpose-built
Arabic tool, from someone who has worked inside the most structurally faithful
Arabic engine anyone shipped, declining to promise that the letters will come
out right. His own essay on why the thing needs to exist names the failure modes
as incorrect letterforms, misplaced dots, broken words — which is where this
piece started.

## What the field is working with

So the sequence runs: the grammar of the script was formally modelled four
decades ago, shipped commercially, won an award — and then in 2021 the one
product that exposed the engine's full range to a designer was withdrawn.
Meanwhile, when a recent paper sets out to *generate* Arabic calligraphy, the
default is to train a GAN on pictures of it: two networks, one producing
candidates and one judging them, converging on output that resembles the
training set. A [systematic review of the
literature](https://thesai.org/Publications/ViewPaper?Volume=16&Issue=3&Code=ijacsa&SerialNo=81)
covering 2009 to 2024 found nineteen studies and reported that GANs and their
variants dominate. Diffusion barely features.

It is worth being precise about how small that is. Of the nineteen, the review's
table of generation approaches lists **four** studies — two on DCGAN, two on
CycleGAN, one on vanilla GAN and VQ-GAN, with a single paper appearing twice
because it tried both. Its table of evaluation methods is
shorter still: four studies assessed by human judgement, and exactly one FID
score in the entire literature. This is not a field with a methodology
problem so much as a field that has barely been staffed.

The same review catalogues the datasets, and they explain a good deal.

<Figure caption="Every dataset the review found, by size, on a log scale: fifteen of them, ten public. Hollow bars are the ones you cannot get. The single large corpus is APTI, forty-five million samples of machine-set printed text, which is not calligraphy and cannot teach a model anything about it. Everything that is actually calligraphic fits in the low thousands. Sizes from Sumayli and Alkaoud (2025), CC BY 4.0.">
  <DatasetLandscape locale="en" />
</Figure>

So the generation literature is a methodological generation behind the rest of
image synthesis, working from a few thousand images, trying to rediscover from
pixels a structure that already exists in a specification.

It is not that nobody digitises this material. IRCICA — the Istanbul institute
that has run the world's triennial calligraphy competition since 1986 — built
its own Ottoman OCR in 2011, because nothing off the shelf could parse the
script. Nearly two million scanned Ottoman pages sit in the Farabi digital
library. [Osmanlica.com](https://www.osmanlica.com/en/) reports average accuracy
above 95%, through a pipeline that reads the page, transliterates the alphabet,
then translates into modern Turkish.

IRCICA holds a calligraphy archive too, including four decades of prize-winning
work from its own competition. The OCR is pointed at the paperwork. Reading the
administration is a solved problem; reading the art is not one anybody has taken
on.

## Where the rules were written, someone encoded them

The tell is what happens next door.

Islamic geometric pattern is having an excellent decade computationally, because
its rules are written down and always were. Square Kufic — *bannā'ī*, masonry
script, the one that looks like pixel art because it was made of bricks — gets
[generated by cellular
automata](https://link.springer.com/article/10.1007/s00004-019-00454-3), which is
almost too neat a fit, since the script genuinely is a grid rule system.

<PostFigure
  src={bannaiTiling}
  alt="Architectural brickwork in the bannā'ī technique, where glazed and unglazed bricks form angular square-Kufic lettering on a wall."
  caption="Bannā'ī brickwork: the letters are the bond. There is no stroke to model here and no pen angle to recover, because the rule is the grid and the grid is visible in the material. A cellular automaton can generate this; nobody needed to train anything. Photograph by Yuriy75, CC BY-SA 3.0."
  locale="en"
/>

Girih gets graph theory. In June, Ugail and Mehmood published a [completion
method with the patterns' own symmetries built in as a
constraint](https://arxiv.org/pdf/2607.02573), scoring edges over a candidate
lattice to reconstruct a full pattern from sparse control geometry — and it does
geometry only, not script.

<PostFigure
  src={girihCompletion}
  alt="A six-panel figure showing an Islamic geometric pattern reconstructed in stages: sparse control points, candidate lattice, rotational orbits, predicted edge scores, symmetry-structured completion, and a rendered gold and blue vector ornament."
  caption="What it looks like when the rules are known: a scatter of control points, a candidate lattice, the rotational orbits the symmetry group permits, scored edges, the completed pattern, and a finished vector ornament. The network never guesses what the pattern should look like. It chooses among the completions the symmetry already permits. I have not found the equivalent diagram for Thuluth. From Ugail and Mehmood (2026), CC BY 4.0."
  locale="en"
/>

Qur'anic page layout is the same story in a different register. The Madani
mushaf's 604 pages at fifteen lines each, with per-line word placement, exist as
[a hand-tuned table](https://qul.tarteel.ai/mushaf_layouts). Nobody trained
anything. Somebody sat down and encoded it, because the layout is a rule and
rules can be written.

And it is not only next door. There is a working market in Arabic calligraphy
software, decades old, and every tool in it that can promise you correct
letterforms gets them the same way: from fonts a type designer drew, assembled
by a program that knows the joining rules. Kelk has been the professional
standard for years. So have eMashq, Kaleam, CalliPro.

<Figure caption="The same split, in a product listing rather than a paper. Four tools hand you letterforms somebody drew and a program assembled; one generates them and, in its own feature list, will only say it preserves the requested text 'up to 95%' and 'reduces' letter and dot mistakes. Style lists as each tool advertises them.">
  <ToolLandscape locale="en" />
</Figure>

And that hedge is, oddly, the honest end of this market. I went through the
other generators — My Qalam AI, the free no-signup sites, the ones that promise
"authentic" Thuluth in seconds — and not one of them says anything about
correctness at all. No accuracy figure, no warning that the letters might come
out wrong, nothing telling you to check the text before you print it. The only
tool that admits it can miss is the one built by somebody who has worked inside
an engine that could not.

The pattern I keep landing on is this. Where the rules are already written down,
somebody encodes them and the results are good. Where they are not, the default
is to train on images and hope. That does not look like a difference in how hard
the art is.

## The part that stayed unwritten

Which leaves the obvious objection. If the proportions were written down in the
tenth century, why is there nothing to train on?

Because the specification and the judgment are two different things, and only
one of them got written. Ibn Muqla's system tells you how tall an alif is. It
does not tell you which of several licensed positions a dot should take in *this*
composition, how far a word may climb before the block goes wrong, or why a
stroke that measures correctly can still be dead on the page. That part moves by
*ijāza*, a licence granted by a master to a student who has demonstrated
mastery, in an unbroken chain back through named teachers. It is not a portfolio
and not a certificate of completion. What travels along that chain is correction:
the master watches the student's hand and adjusts it. The judgment lives in the
correction, and the correction was never written down, because it never needed to
be. There was always a person.

<PostFigure
  src={ijazaMansour}
  alt="An illuminated ijāza certificate on deep blue with gold floral borders: a line of Thuluth at the top, a block of Naskh beneath it, and two small cartouches of dense script at the bottom carrying the grant and its date."
  caption="An ijāza belonging to the calligrapher Nassar Mansour, who trained under Hasan Çelebi in Istanbul and in 2003 became the first Jordanian to receive the traditional licence. It is permission to sign your work, and is itself a piece of calligraphy: the demonstration above, the grant and the chain of teachers in the cartouches below. What it certifies is not that the holder knows the rules, but that a named master watched them write. Photograph by Asilmansur7, CC BY-SA 4.0."
  locale="en"
/>

That is a strength of the tradition, and it is also why there is no corpus.
Centuries of accumulated judgment about pen angle, about where a dot may move and
how far, about which letters may be stacked and in what order — held in hands,
taught in rooms, and almost entirely absent from any machine-readable form.

Whether you could grant a machine an ijāza is a fun question and not the
important one. The important one is that this half was never written down, so
there is nothing to train on but the finished work — and the finished work is the
one part that does not contain it.

## What might actually help

Not more images. Strokes.

[Calliar](https://arxiv.org/abs/2106.10745), from the ARBML community, is the
only dataset I know of that treats Arabic calligraphy as motion: 2,500 sentences
traced with a digital pen by four annotators, recorded as ordered point
sequences and annotated at stroke, character, word and sentence level. Every
other dataset in this field hands the model a picture of dried ink. Calliar
hands it the gesture.

<Figure caption="Calliar at each level, with the training split filled in. Strokes are the largest count in the set and the level nobody has trained on: 45,572 of them, 36,561 available to train. A stroke is one uninterrupted movement of the pen, and a dot counts as a stroke of its own. The printed-text corpus in the chart above holds 45.31 million samples. Numbers from Calliar Table 4 (Alyafeai et al., 2021), CC BY 4.0.">
  <CalliarComposition locale="en" />
</Figure>

<PostFigure
  src={calliarStrokes}
  alt="The JSON stroke representation of the Arabic word khalid, five colour-coded lists of coordinate pairs, shown beside the rendered word with each stroke drawn in its matching colour."
  caption="One word, خالد, as Calliar stores it: five strokes, each a labelled list of points in the order the hand made them. The dot is its own stroke. This is the only representation in the field that records writing as something that happened rather than something that exists. From Calliar (Alyafeai et al., 2021), CC BY 4.0."
  locale="en"
/>

<PostFigure
  src={calliarSamples}
  alt="A four-by-four grid of the basmala written sixteen different ways, each rendered as coloured strokes on white, showing wide variation in composition and proportion."
  caption="Sixteen writings of the same phrase, each stroke in its own colour. The variation between them is not noise around a correct answer. It is the range a trained hand is licensed to move in — the thing a model would have to learn, and the thing no collection of finished images contains. From Calliar (Alyafeai et al., 2021), CC BY 4.0."
  locale="en"
/>

There is a mature technique waiting for exactly this shape of data. Graves showed
in [2013](https://arxiv.org/abs/1308.0850) that an LSTM predicting pen
trajectories through a Gaussian mixture can generate convincing handwriting, and can be primed with a real writer's pen
movements to write in that person's hand. The lineage runs through to transformer
ink models today. Nobody, as far as I can find, has run that playbook on Calliar.
Thirty-six thousand training strokes is probably too few. It is the right shape
of too few.

The physical end is emptier still. Chinese calligraphy has parameterised brush
models — 3D hair geometry, ink deposition along the trajectory — and [robot arms
driven by optimal control](https://arxiv.org/abs/1911.08002) against those
models. The qalam is a reed cut at a
fixed angle: a rigid nib with a straight contact edge, which is what produces the
entire thick-and-thin logic of every classical style. It is a *simpler* object to
model than a soft brush, and its geometry is the whole grammar of the stroke. I
could not find anyone who has modelled it.

The closest thing is a project rather than a system. In a 2021–22 seminar at the
Berlin University of the Arts, Salam Shokor taught a UR5 industrial arm to write
Arabic, using a 3D-printed holder, a broad-nib parallel pen standing in for the
reed, and a custom stroke-based Arabic font built in Grasshopper. The font was
explicitly inspired by Hofstadter's *Letter Spirit*, the 1990s project that tried
to generate a whole alphabet in a consistent style from a couple of seed letters.
It worked on a grid of fifty-six segments rather than a canvas of pixels, because
its authors thought the hard part was holding two kinds of sameness at once: what
makes an *a* an *a*, and what makes it Helvetica. Which is to say the robot
drives from a font, along paths a person defined. Even the robot is running on
written-down rules. Nobody is simulating the nib.

<PostFigure
  src={qalamCarving}
  alt="A knife carving the nib of a reed calligraphy pen on a wooden cutting block, with shavings scattered and two split reed nibs showing their hollow interiors alongside."
  caption="A qalam being cut. Everything a Chinese brush model has to solve with hair dynamics and ink flow, this tool settles with one angled facet and a slit — which is why the same stroke is thick descending and thin across, in every classical style, without the calligrapher doing anything but turning. It is the most modellable instrument in the whole discipline and nobody has modelled it. Photograph by Faizal Somadi, CC BY-SA 4.0."
  locale="en"
/>

It is worth being exact about how absent Arabic is from this. Visual text
rendering *has* been largely solved: [Glyph-ByT5-v2](https://arxiv.org/abs/2406.10208)
renders accurate text in ten languages and ships a benchmark to prove it. The
ten are English, Chinese, French, German, Spanish, Portuguese, Italian, Russian,
Japanese and Korean. Arabic is in neither the model nor the benchmark, so there
is not even a number to be bad at.

Meanwhile the region is building the infrastructure and hitting the same wall.
Qatar's [Fanar 2.0](https://huggingface.co/QCRI/Fanar-2-Oryx-IG) shipped in
March, a genuinely serious Arabic-first stack, with an image model built on FLUX and
trained on 480,000 culturally-curated images gathered across twenty-two Arab
countries. It scores 85.49 on the card's own cultural-compliance
measure. And the same card says, in plain words, that text rendering in images —
especially Arabic — remains challenging. A purpose-built Arabic image model that cannot write Arabic. The
companion vision model *recognises* calligraphy as a listed feature. Reading and
writing have come apart.

## What it is good for today

Vectorisation, and not much else.

Turning a scanned piece into curves has always meant auto-tracing it and then
repairing the places where the tracer turned a confident sweep into a dozen
anchor points and a wobble. Learned vectorisers are better at the thing that
actually matters here, which is not coverage but curve continuity: the
smoothness of a single stroke, which is the first thing a trained eye checks and
close to the last thing a threshold-and-trace pipeline preserves. A [2025 method
reports near-perfect overlap and smoother curvature than standard auto-tracing
across Naskh, Ruq'ah and Diwani](https://www.academia.edu/145562424/A_Novel_Deep_Learning_Approach_for_High_Fidelity_Vectorization_of_Arabic_Calligraphy).
I would treat the exact figures with care given where it is published, but the
direction is unsurprising.

Note what that task requires: nothing. The model does not need to know Arabic,
or which style it is looking at, or that the marks are writing at all. It is
tracing curves. That is the honest shape of what machine learning currently
offers this craft — real, useful, and entirely outside the part that is hard.

The rest of it is worth saying plainly, because the marketplaces are now full of
generated Qur'anic verses sold as wall art, and a mangled verse is not a
typographical infelicity. If the text is sacred, take it from a verified mushaf
and set it with a font that knows the script. The generator does not know what
it wrote.

## What I could not settle

None of this is an argument that the thing cannot be done. It is an argument,
from someone who is neither a researcher in this field nor a master of the
craft, that it is being approached from the wrong end — and I would rather be
shown wrong about that than right.

The structure of the script is known: DecoType demonstrated that decades ago,
and their model is documented in public even if the product is gone. The
calligraphic rules on top of it are known too; they are just held in people
rather than files. The datasets that would matter are the ones recording the
hand in motion, and one of them exists and is small and open and largely unused.

What is missing is not compute, or model capacity, or volume — this was never
going to be fixed by more of what already exists. It is a different kind of
record altogether, one nobody has made, of a body of knowledge whose entire
transmission mechanism was designed around not needing one. Tarteel got 67 hours of Qur'anic recitation from twelve
hundred volunteers in six months, which suggests the community can build the
corpus when someone frames the task. For calligraphy, nobody has framed it.

Several things I could not settle, and would take a correction on. Whether
anyone has run the stroke-generation playbook on Calliar and simply not
published it. How much of DecoType's model is documented well enough to rebuild
from, now that the product is gone. Whether there is Arabic robotic calligraphy
work in a language I do not read: the literature on brush robots is largely
Chinese, and I searched in English and Arabic only. And how far the disputed
figures go: sources do not even agree on how many dots high a Thuluth alif is,
which is a strange place for a thousand-year-old specification to be vague.

What I am fairly confident of is the shape of it. The window is not open
indefinitely: the chain that carries this knowledge runs through living people,
and what they hold is the correction — the adjustment made to a hand in a room,
which no photograph of a finished panel has ever contained. Every archive being
built right now is an archive of dried ink.

The thing worth recording is the stroke.
