Article29 min read

Why AI gets Arabic calligraphy wrong

One model reads Arabic handwriting at under one per cent word error. The best anyone has managed on classical calligraphy is forty-five. I went looking for the missing data that would explain the gap and found the opposite — a script whose rules were written down in the tenth century, and again in software in the 1980s.

Type a line of Arabic into any of the calligraphy generators and you get back something with the correct temperature and no content. The strokes have weight. The composition stacks the way Thuluth stacks. Somebody who has never read a word of Arabic would accept it without a second look, and that is exactly the population it was built for.

Anyone who reads Arabic sees it in under a second. Letters that should have joined are sitting apart. A ه has taken its initial form in the middle of a word. The dots have drifted off the letters they belong to and settled wherever the density looked right. It is not bad calligraphy. Bad calligraphy is a thing a beginner produces, and it is legible. This is calligraphy-shaped noise.

I wanted to know why. I build software for a living and I have spent enough years failing to write a decent alif to know roughly what the craft asks of you, which is a strange pair of interests to have but a useful one for this particular question. What follows is what I found looking into it, including the part where the explanation I assumed was true turned out to be wrong.

Thuluth by Yaqut al-Musta'simi, the thirteenth-century master whose hand is still a reference seven centuries on. Every relationship in this panel — how far a stroke travels before it turns, where a dot sits relative to the letter that owns it, which words climb and which stay on the line — is a decision, and the ones the generators get wrong are all here. Library of Congress, public domain.

Not for want of data

My first assumption was the obvious one: Arabic is under-resourced, and the models will catch up. That was true for a long time. It has stopped being true, and the numbers now say something stranger.

Qalam, a vision encoder-decoder trained on four and a half million manuscript images, reads Arabic handwriting at 0.80% word error. QARI-OCR, a 2B-parameter Qwen2-VL fine-tune, reads diacritically-rich Arabic print at 6.1% character error — with the tashkeel, which is the hard part. Both are open. Both run on hardware you own.

Three numbers recur below, so it is worth fixing them now. Word error rate is the share of words a system gets wrong; character error rate the share of characters, which is the gentler measure since a word can be wrong by one letter. Exact match is the strict one: the share of images transcribed with nothing wrong at all.

Then somebody built a benchmark for calligraphy specifically. DuwatBench landed in January, out of MBZUAI: 1,272 samples across Thuluth, Diwani, Kufic, Naskh, Ruq’ah and Nasta’liq, with transcriptions, style labels and bounding boxes. Thirteen models were pointed at it, open and closed, general-purpose and Arabic-specific.

One phrase — بسم الله الرحمن الرحيم — in each of the six styles the benchmark covers, with how many samples of each it holds. These are the same four words every time. A model that has learned Naskh has not learned Thuluth, and the sample counts show which end of that range the data sits at. From DuwatBench (Patle et al., 2026), CC BY 4.0.

The best result on the board is Gemini 2.5 Flash, at 45% word error and 42% exact match. Across the thirteen models tested, exact match runs from zero to that 42%, and averages 17.9%. The open models are worse in ways that are almost funny to read about: one of them responds to difficulty by emitting the word Allah over and over, having correctly learned that this is a good bet on Arabic calligraphy and nothing else about the image in front of it.

Eleven images across the top, the correct reading beneath them, and then every model's attempt. Green is a hit. Read down the InternVL row and watch it answer الله to almost everything; read the TrOCR row and watch it return a single dot. GPT-4o's contribution to several columns is a refusal — it declines to help, in Arabic. The models are not slightly wrong here. They are not reading. From DuwatBench (Patle et al., 2026), CC BY 4.0.

So a system built for Arabic manuscripts reads them at under one per cent word error, and the best system anyone has pointed at a calligraphic panel gets 45%. Nor is this general-purpose models being asked something outside their brief. Two of the thirteen were Arabic-specific, and both sit low on the board: Microsoft’s Arabic handwriting TrOCR returns 99.98% word error and not one exact match across the entire benchmark. A model built for Arabic handwriting, handed Arabic handwriting, reading essentially nothing.

I expected the per-style breakdown to be a clean ramp: plain scripts easy, ornate scripts hard. It is not, and the way it fails to be is the most interesting thing in the paper.

Word error rate by calligraphic styleA bar chart of word error rate for six Arabic calligraphic styles, comparing the best closed-source model, Gemini 2.5 Flash, with the best open-source model, Gemma 3 27B. Thuluth 35 versus 63 per cent, Naskh 48 versus 51, Nasta'liq 52 versus 66, Diwani 57 versus 73, Ruq'ah 58 versus 76, and Kufic 71 versus 78. Lower is better.Gemini 2.5 FlashGemma 3 27Bword error rate — lower is betterThuluth35%63%Naskh48%51%Nasta'liq52%66%Diwani57%73%Ruq'ah58%76%Kufic71%78%
Word error rate for each style, best closed model against best open one. Not a ramp. Kufic is worst for both, Naskh is the steadiest, and Thuluth — the ornate one — is where the closed model does best of all six, twenty-eight points clear of the open model on the same images. Numbers from DuwatBench Table 3 (Patle et al., 2026), CC BY 4.0.

Kufic being the hardest is not obvious. It is the most geometric of the six, the one closest to something you could describe with a grid, and it defeats everything, presumably because that same geometry strips out the redundancy a reader normally leans on.

The Thuluth result I do not think is a reading result at all. Thuluth is 55% of this benchmark and the corpus is 44% Qur’anic, so a model with a strong language prior has a great deal of it effectively memorised, and guessing the basmala is not the same skill as reading it. This looks at first like a contradiction with the paper’s own summary, which says Naskh and Ruq’ah are the most accurately recognised styles. It is not. That holds across all thirteen models averaged together, and the same passage notes separately that Thuluth and Diwani run substantially higher for the closed-source models in particular. The Gemini column is the second half of that sentence.

Not for want of scale

The comfortable next move is to say this is still a data problem, just a bigger one. A paper from June puts that to rest.

Sana Al-azzawi, Elisa Barney and Marcus Liwicki ran one model across nine handwriting datasets — Arabic, Urdu, Persian, English, German — holding the architecture fixed and varying only the script and the volume of training data. The Arabic-script penalty is largest when data is scarce, narrows as data grows, and then stops narrowing at a consistent five to seven character-error points. Cleaning the annotations helps and does not close it.

Their diagnosis is the thing you notice within about ten minutes of picking up a pen. Around 30% of the substitution errors on Arabic-script data are confusions between visually similar characters, against about 15% on Latin. One group of five letters in this alphabet is told apart by nothing but dots. Standing alone they differ a little; inside a word they collapse onto the same stroke and the dots are the whole of what is left. In a Latin alphabet the letterform carries the identity and the diacritic modifies it. Here the diacritic is the identity, riding on a stroke that means nothing by itself.

Five Arabic letters that differ only by their dotsThe top row repeats a single dotless stroke five times: the shared medial skeleton of five Arabic letters. The bottom row shows the same five letters in their medial forms with their dots — baa with one below, taa with two above, thaa with three above, noon with one above, and yaa with two below — which are the only marks distinguishing them.the shared skeleton, mid-wordwhat tells them apart‍ب‍‍ت‍‍ث‍‍ن‍‍ي‍
Five letters in the middle of a word. There is no character for the top row — it had to be drawn, because no font will give you a letter with its dots taken off. That is the point: the shared stroke is not a letter, it is the part five letters have in common, and everything that makes it a particular word is the marks above and below it.

Now put that in a Thuluth composition, where the calligrapher moves dots off their letters to balance the block — a licensed, traditional, expected move that every trained reader parses without effort because they are reading the word, not the glyph. The model is reading glyphs. It was always reading glyphs.

A Qur'an folio in early Kufic, eighth or ninth century. Look closely at what the dots are doing: the letters carry no pointing at all — no i'jam, nothing to tell a bāʾ from a tāʾ — while the red dots scattered across the page are a separate, older system marking vowels. Arabic acquired its dots in layers, for different jobs, centuries apart. Public domain.

If you want the number that makes this unarguable, look at Ajami, the Arabic-script orthographies of West Africa. The first HTR dataset for Fulfulde and Hausa manuscripts was published in the ICDAR 2025 proceedings. Its authors ran a suite of Arabic-script recognition models built for historical manuscripts over it, and report character error rates of 65 to 84%. Not word error — character error, the gentler of the two measures. Up to five letters in six wrong, one letter at a time, by models built for exactly this kind of material. “Arabic AI” turns out to mean a narrow slice of Arabic.

A page from a poem by Abdullahi dan Fodio, of the Sokoto Caliphate. Everything here is Arabic script and almost none of it is what a model trained on Cairo has seen: a West African hand, a second smaller hand crowding the glosses, its own orthographic conventions, a decorated panel that is not text at all. Before a model can misread this it has to work out which marks are even the line. British Library / Endangered Archives Programme, CC0.

What the script is doing

At this point I stopped reading benchmark papers and went back to the script itself, because the errors had started to look structural rather than statistical.

Arabic is not a set of glyphs you substitute one for one. A letter takes a different form depending on whether it begins a word, sits inside it, ends it, or stands alone. Six letters — ا د ذ ر ز و — refuse to join to whatever follows them, so a single word breaks into several connected runs. Certain pairs fuse into ligatures that are not either letter. In the classical styles the line itself stops being horizontal: Diwani sweeps, Thuluth stacks words vertically into a block, and where a word sits is a compositional decision made by a person weighing the whole panel.

Why substituting one glyph for another cannot produce ArabicTwo rows. The first shows the single Arabic letter ayn in its four contextual forms — isolated, initial, medial and final — each a different shape selected by the letter's position in a word. The second shows the word madrasa broken into three separately joined runs — meem-dal, ra, and seen-ta marbuta — because dal and ra do not join to whatever follows them.one letter, four shapesعisolatedع‍initial‍ع‍medial‍عfinalone word, three runsمدرسةمدرسة
The two rules that a character-for-character swap cannot honour. Above: one letter, four shapes, chosen by its neighbours. Below: a word that is not one connected run but three, because dāl and rāʾ refuse to join to what follows them. Those gaps are not spaces, whatever a renderer might assume. Get either wrong and a native reader sees it immediately.

None of that is decoration on top of the letters. It is the letters. The script has a grammar, and the grammar is the art.

The alphabet set out in Thuluth by 'Ala' al-Din Tabrizi. This is the exercise every student copies, and it is also, read the right way, a specification: the same letter appears in its different positions, the proportions are fixed, and what looks like a display of virtuosity is a table of rules. Dated 1593; Library of the Golestan Palace, via Reed College. Public domain.

And the line does not have to be a line. Diwani, the Ottoman chancery hand, packs its words into a rising sweep and closes the gaps until the block reads as one mass, a habit traditionally explained as a security measure, since a text with no room between its words is a text nobody can insert a clause into.

An imperial decree of Mehmed II in Diwani. The words climb, the spacing is deliberately eliminated, and the whole block is engineered to be unforgeable and unamendable. Every property that made it good chancery practice makes it hostile to a model: no baseline, no word gaps, no isolated forms to latch onto. Google Art Project, public domain.

Nasta’liq goes the other way. The name contracts naskh and ta’liq, and ta’liq means suspension, which is what the script does: words hang off one another and descend as the line runs, so a reader tracks a diagonal rather than a baseline. A layout engine that assumes text sits on a horizontal has already lost.

A Nasta'liq panel. There is no baseline here in any sense a text layout engine would recognise — words hang off each other and the eye follows the slope down and to the left. This is the style DuwatBench holds the fewest samples of — sixty-seven. Public domain.

A model trained on images of finished calligraphy is being asked to recover that grammar from its residue. It sees where the ink landed. It never sees the rule that put it there, or the pen angle, or the order the strokes were made in, or the fact that the calligrapher moved that dot deliberately. Given enough examples it learns the texture of the residue, which is precisely what it then produces.

The rules were already written down

Here is where I expected to find the gap — some vagueness in the tradition, a body of knowledge too intuitive to commit to paper — and found close to the opposite.

The proportions are not taste passed down as vibes. They are a ratio system, and it was formalised more than a thousand years ago as al-khaṭṭ al-manṣūb, proportioned script, traditionally credited to the Abbasid vizier and calligrapher Ibn Muqla, who died in 940. It has three units.

The proportional system of Arabic calligraphyThree panels. First, the rhombic dot, produced by pressing the reed nib once at the angle it was cut to. Second, an alif whose height is measured as a stack of such dots — five of them in Naskh, with the count fixed for each style. Third, a circle whose diameter equals the height of the alif. Every letter in the classical styles is proportioned against these three units, a system formalised as proportioned script and traditionally credited to Ibn Muqla.the rhombic dotone press of the nib,held at its cut angle5 dotsthe aliffive dots in Naskh —the count is fixed per stylethe circleits diameter isthe alif
The dot is not punctuation here: it is the ruler. Press the cut nib once and you get a rhombus; stack a fixed number of them and you have the height of an alif, five of them in Naskh and a different count for each style; take that height as a diameter and you have the circle every curve is checked against.

That system is why a trained reader can look at a letter and say it is incorrect rather than just unfamiliar: there is a number it failed to meet. Ibn Muqla is credited with canonising the six classical pens on the same basis, and Ibn al-Bawwab and Yaqut al-Musta’simi refined the system after him. A thousand years before anyone wrote a font engine, this script had a written specification.

And it did not stop at the letter. Whole page layouts were fixed forms too, reproduced by generations of calligraphers to the same plan.

A hilye by Hafiz Osman (1642–1698), who devised the layout that every later hilye follows. This is a template: the basmala across the top, the round belly at the centre carrying Ali ibn Abi Talib's description of the Prophet, the four corner medallions naming the first caliphs, the verse band beneath. All in fixed relation, and every hilye written since sits inside those constraints. Sadberk Hanım Museum, public domain.

That is the tenth century. The modern version is the part that actually reframed this for me.

DecoType was incorporated in 1985 by Peter Somers, Mirjam Somers and Thomas Milo — Designers of Computer-aided Typography. What they built was not a glyph set but a font that knew the script’s rules and applied them: their Ruq’ah was the first dynamic Arabic typeface of its kind, and by the early 1990s it was shipping inside Microsoft Office. The DecoType Naskh and Thuluth families went out to a generation of Office users under an “MS Office Edition” label, on licences dated 1992 to 1998. Whether current builds still carry them I could not establish.

Underneath sat ACE, the Arabic Calligraphic Engine, built first for Ruq’ah and then extended across a broad analysis of Naskh, and later renamed the Advanced Composition Engine as it grew to drive any Arabic typeface.

Milo’s own account of the model is explicit about the stakes: ACE approaches Arabic typography in analogy with pre-typographic text manufacture. It encodes the orthographic rules by which letterform combinations weave into ligature forms. It is a considered rejection of movable type as a model for this script, a rejection made by people who went and studied the manuscript tradition first. He took the Peter Karow Award for it in 2009.

It reached working designers as Tasmeem, the Arabic plug-in for InDesign ME that DecoType built with WinSoft International, who published it. As far as I can tell it remains the most structurally faithful Arabic composition system anyone has shipped.

WinSoft discontinued it in April 2021 — the product, not DecoType, whose engine had always been the part underneath. Last release for InDesign CC 2021. Existing clients supported to the end of their subscriptions, no new sales, no further upgrades. DecoType’s own site now reads as an archive; the most recent thing dated on it is a 2017 lecture.

One of the people who worked on that plug-in is still at it, and what he sells now is the neatest summary of where this has ended up. Nihad Taisir Nadam — Damascus-born, Dubai-based — co-developed the first online Arabic digital calligraphy tool with WinSoft in 2011, alongside Tasmeem itself, and by his own account has worked with Adobe, WinSoft and DecoType on Arabic font technology. He now sells ArabicDesign.ai, which generates calligraphy from typed Arabic in about fifteen styles and exports SVG.

Read its feature list closely. It offers to preserve the requested text at up to 95%, and to reduce common letter and dot mistakes. That is a purpose-built Arabic tool, from someone who has worked inside the most structurally faithful Arabic engine anyone shipped, declining to promise that the letters will come out right. His own essay on why the thing needs to exist names the failure modes as incorrect letterforms, misplaced dots, broken words — which is where this piece started.

What the field is working with

So the sequence runs: the grammar of the script was formally modelled four decades ago, shipped commercially, won an award — and then in 2021 the one product that exposed the engine’s full range to a designer was withdrawn. Meanwhile, when a recent paper sets out to generate Arabic calligraphy, the default is to train a GAN on pictures of it: two networks, one producing candidates and one judging them, converging on output that resembles the training set. A systematic review of the literature covering 2009 to 2024 found nineteen studies and reported that GANs and their variants dominate. Diffusion barely features.

It is worth being precise about how small that is. Of the nineteen, the review’s table of generation approaches lists four studies — two on DCGAN, two on CycleGAN, one on vanilla GAN and VQ-GAN, with a single paper appearing twice because it tried both. Its table of evaluation methods is shorter still: four studies assessed by human judgement, and exactly one FID score in the entire literature. This is not a field with a methodology problem so much as a field that has barely been staffed.

The same review catalogues the datasets, and they explain a good deal.

Arabic calligraphy datasets by sizeA logarithmic bar chart of ten Arabic script datasets by number of samples. APTI leads with 45.31 million samples but is machine-set printed text rather than calligraphy. KHATT has 9,327; Alrehali et al. 5,240, private; HICMA 5,031; ACL 3,467; Khayyat et al. 2,653, private; Calliar 2,500; Kaoudja et al. 1,685; Belila and Gasmi 900; and Allaf et al. 267, private.publicprivatesamples — log scaleAPTI45.31MKHATT9,327Alrehali et al.5,240HICMA5,031ACL3,467Khayyat et al.2,653Calliar2,500Kaoudja et al.1,685Belila & Gasmi900Allaf et al.267
Every dataset the review found, by size, on a log scale: fifteen of them, ten public. Hollow bars are the ones you cannot get. The single large corpus is APTI, forty-five million samples of machine-set printed text, which is not calligraphy and cannot teach a model anything about it. Everything that is actually calligraphic fits in the low thousands. Sizes from Sumayli and Alkaoud (2025), CC BY 4.0.

So the generation literature is a methodological generation behind the rest of image synthesis, working from a few thousand images, trying to rediscover from pixels a structure that already exists in a specification.

It is not that nobody digitises this material. IRCICA — the Istanbul institute that has run the world’s triennial calligraphy competition since 1986 — built its own Ottoman OCR in 2011, because nothing off the shelf could parse the script. Nearly two million scanned Ottoman pages sit in the Farabi digital library. Osmanlica.com reports average accuracy above 95%, through a pipeline that reads the page, transliterates the alphabet, then translates into modern Turkish.

IRCICA holds a calligraphy archive too, including four decades of prize-winning work from its own competition. The OCR is pointed at the paperwork. Reading the administration is a solved problem; reading the art is not one anybody has taken on.

Where the rules were written, someone encoded them

The tell is what happens next door.

Islamic geometric pattern is having an excellent decade computationally, because its rules are written down and always were. Square Kufic — bannā’ī, masonry script, the one that looks like pixel art because it was made of bricks — gets generated by cellular automata, which is almost too neat a fit, since the script genuinely is a grid rule system.

Bannā'ī brickwork: the letters are the bond. There is no stroke to model here and no pen angle to recover, because the rule is the grid and the grid is visible in the material. A cellular automaton can generate this; nobody needed to train anything. Photograph by Yuriy75, CC BY-SA 3.0.

Girih gets graph theory. In June, Ugail and Mehmood published a completion method with the patterns’ own symmetries built in as a constraint, scoring edges over a candidate lattice to reconstruct a full pattern from sparse control geometry — and it does geometry only, not script.

What it looks like when the rules are known: a scatter of control points, a candidate lattice, the rotational orbits the symmetry group permits, scored edges, the completed pattern, and a finished vector ornament. The network never guesses what the pattern should look like. It chooses among the completions the symmetry already permits. I have not found the equivalent diagram for Thuluth. From Ugail and Mehmood (2026), CC BY 4.0.

Qur’anic page layout is the same story in a different register. The Madani mushaf’s 604 pages at fifteen lines each, with per-line word placement, exist as a hand-tuned table. Nobody trained anything. Somebody sat down and encoded it, because the layout is a rule and rules can be written.

And it is not only next door. There is a working market in Arabic calligraphy software, decades old, and every tool in it that can promise you correct letterforms gets them the same way: from fonts a type designer drew, assembled by a program that knows the joining rules. Kelk has been the professional standard for years. So have eMashq, Kaleam, CalliPro.

Tools that encode the script

Kelk
Naskh, Nasta'liq, Thuluth, Diwani khafi and jali, Kufi, Ruq'ah, Shekasteh, Moalla
Drawn by type designers, assembled by the program
eMashq
Thuluth, Ijazah, Diwani, Diwani jali, Naskh, Nasta'liq, Shikasta, Ruq'ah
Drawn by type designers, assembled by the program
Kaleam
Thuluth, Diwani, Persian and other traditional styles
Drawn by type designers, assembled by the program
CalliPro
21 fonts — Diwan Naskh Mishafi, Diwan Thuluth, Waseem, Kufi
Drawn by type designers, assembled by the program

Tools that generate the shapes

ArabicDesign.ai
About fifteen styles, from a typed prompt, exports SVG
“Preserves the requested text at up to 95%”, and “reduces” common letter and dot mistakes

Style lists as each tool advertises them. The last row is quoted from its own feature list — the only one of the five that does not simply hand you the letters.

The same split, in a product listing rather than a paper. Four tools hand you letterforms somebody drew and a program assembled; one generates them and, in its own feature list, will only say it preserves the requested text 'up to 95%' and 'reduces' letter and dot mistakes. Style lists as each tool advertises them.

And that hedge is, oddly, the honest end of this market. I went through the other generators — My Qalam AI, the free no-signup sites, the ones that promise “authentic” Thuluth in seconds — and not one of them says anything about correctness at all. No accuracy figure, no warning that the letters might come out wrong, nothing telling you to check the text before you print it. The only tool that admits it can miss is the one built by somebody who has worked inside an engine that could not.

The pattern I keep landing on is this. Where the rules are already written down, somebody encodes them and the results are good. Where they are not, the default is to train on images and hope. That does not look like a difference in how hard the art is.

The part that stayed unwritten

Which leaves the obvious objection. If the proportions were written down in the tenth century, why is there nothing to train on?

Because the specification and the judgment are two different things, and only one of them got written. Ibn Muqla’s system tells you how tall an alif is. It does not tell you which of several licensed positions a dot should take in this composition, how far a word may climb before the block goes wrong, or why a stroke that measures correctly can still be dead on the page. That part moves by ijāza, a licence granted by a master to a student who has demonstrated mastery, in an unbroken chain back through named teachers. It is not a portfolio and not a certificate of completion. What travels along that chain is correction: the master watches the student’s hand and adjusts it. The judgment lives in the correction, and the correction was never written down, because it never needed to be. There was always a person.

An ijāza belonging to the calligrapher Nassar Mansour, who trained under Hasan Çelebi in Istanbul and in 2003 became the first Jordanian to receive the traditional licence. It is permission to sign your work, and is itself a piece of calligraphy: the demonstration above, the grant and the chain of teachers in the cartouches below. What it certifies is not that the holder knows the rules, but that a named master watched them write. Photograph by Asilmansur7, CC BY-SA 4.0.

That is a strength of the tradition, and it is also why there is no corpus. Centuries of accumulated judgment about pen angle, about where a dot may move and how far, about which letters may be stacked and in what order — held in hands, taught in rooms, and almost entirely absent from any machine-readable form.

Whether you could grant a machine an ijāza is a fun question and not the important one. The important one is that this half was never written down, so there is nothing to train on but the finished work — and the finished work is the one part that does not contain it.

What might actually help

Not more images. Strokes.

Calliar, from the ARBML community, is the only dataset I know of that treats Arabic calligraphy as motion: 2,500 sentences traced with a digital pen by four annotators, recorded as ordered point sequences and annotated at stroke, character, word and sentence level. Every other dataset in this field hands the model a picture of dried ink. Calliar hands it the gesture.

What Calliar contains, by annotation levelA bar chart of the Calliar dataset at four annotation levels, showing the training split against the held-out portion. Strokes: 36,561 of 45,572. Characters: 24,722 of 30,720. Words: 6,065 of 7,556. Sentences: 2,000 of 2,500.training splitvalidation + testthe entire datasetstrokes⁦36,561 / 45,572⁩characters⁦24,722 / 30,720⁩words⁦6,065 / 7,556⁩sentences⁦2,000 / 2,500⁩
Calliar at each level, with the training split filled in. Strokes are the largest count in the set and the level nobody has trained on: 45,572 of them, 36,561 available to train. A stroke is one uninterrupted movement of the pen, and a dot counts as a stroke of its own. The printed-text corpus in the chart above holds 45.31 million samples. Numbers from Calliar Table 4 (Alyafeai et al., 2021), CC BY 4.0.
One word, خالد, as Calliar stores it: five strokes, each a labelled list of points in the order the hand made them. The dot is its own stroke. This is the only representation in the field that records writing as something that happened rather than something that exists. From Calliar (Alyafeai et al., 2021), CC BY 4.0.
Sixteen writings of the same phrase, each stroke in its own colour. The variation between them is not noise around a correct answer. It is the range a trained hand is licensed to move in — the thing a model would have to learn, and the thing no collection of finished images contains. From Calliar (Alyafeai et al., 2021), CC BY 4.0.

There is a mature technique waiting for exactly this shape of data. Graves showed in 2013 that an LSTM predicting pen trajectories through a Gaussian mixture can generate convincing handwriting, and can be primed with a real writer’s pen movements to write in that person’s hand. The lineage runs through to transformer ink models today. Nobody, as far as I can find, has run that playbook on Calliar. Thirty-six thousand training strokes is probably too few. It is the right shape of too few.

The physical end is emptier still. Chinese calligraphy has parameterised brush models — 3D hair geometry, ink deposition along the trajectory — and robot arms driven by optimal control against those models. The qalam is a reed cut at a fixed angle: a rigid nib with a straight contact edge, which is what produces the entire thick-and-thin logic of every classical style. It is a simpler object to model than a soft brush, and its geometry is the whole grammar of the stroke. I could not find anyone who has modelled it.

The closest thing is a project rather than a system. In a 2021–22 seminar at the Berlin University of the Arts, Salam Shokor taught a UR5 industrial arm to write Arabic, using a 3D-printed holder, a broad-nib parallel pen standing in for the reed, and a custom stroke-based Arabic font built in Grasshopper. The font was explicitly inspired by Hofstadter’s Letter Spirit, the 1990s project that tried to generate a whole alphabet in a consistent style from a couple of seed letters. It worked on a grid of fifty-six segments rather than a canvas of pixels, because its authors thought the hard part was holding two kinds of sameness at once: what makes an a an a, and what makes it Helvetica. Which is to say the robot drives from a font, along paths a person defined. Even the robot is running on written-down rules. Nobody is simulating the nib.

A qalam being cut. Everything a Chinese brush model has to solve with hair dynamics and ink flow, this tool settles with one angled facet and a slit — which is why the same stroke is thick descending and thin across, in every classical style, without the calligrapher doing anything but turning. It is the most modellable instrument in the whole discipline and nobody has modelled it. Photograph by Faizal Somadi, CC BY-SA 4.0.

It is worth being exact about how absent Arabic is from this. Visual text rendering has been largely solved: Glyph-ByT5-v2 renders accurate text in ten languages and ships a benchmark to prove it. The ten are English, Chinese, French, German, Spanish, Portuguese, Italian, Russian, Japanese and Korean. Arabic is in neither the model nor the benchmark, so there is not even a number to be bad at.

Meanwhile the region is building the infrastructure and hitting the same wall. Qatar’s Fanar 2.0 shipped in March, a genuinely serious Arabic-first stack, with an image model built on FLUX and trained on 480,000 culturally-curated images gathered across twenty-two Arab countries. It scores 85.49 on the card’s own cultural-compliance measure. And the same card says, in plain words, that text rendering in images — especially Arabic — remains challenging. A purpose-built Arabic image model that cannot write Arabic. The companion vision model recognises calligraphy as a listed feature. Reading and writing have come apart.

What it is good for today

Vectorisation, and not much else.

Turning a scanned piece into curves has always meant auto-tracing it and then repairing the places where the tracer turned a confident sweep into a dozen anchor points and a wobble. Learned vectorisers are better at the thing that actually matters here, which is not coverage but curve continuity: the smoothness of a single stroke, which is the first thing a trained eye checks and close to the last thing a threshold-and-trace pipeline preserves. A 2025 method reports near-perfect overlap and smoother curvature than standard auto-tracing across Naskh, Ruq’ah and Diwani. I would treat the exact figures with care given where it is published, but the direction is unsurprising.

Note what that task requires: nothing. The model does not need to know Arabic, or which style it is looking at, or that the marks are writing at all. It is tracing curves. That is the honest shape of what machine learning currently offers this craft — real, useful, and entirely outside the part that is hard.

The rest of it is worth saying plainly, because the marketplaces are now full of generated Qur’anic verses sold as wall art, and a mangled verse is not a typographical infelicity. If the text is sacred, take it from a verified mushaf and set it with a font that knows the script. The generator does not know what it wrote.

What I could not settle

None of this is an argument that the thing cannot be done. It is an argument, from someone who is neither a researcher in this field nor a master of the craft, that it is being approached from the wrong end — and I would rather be shown wrong about that than right.

The structure of the script is known: DecoType demonstrated that decades ago, and their model is documented in public even if the product is gone. The calligraphic rules on top of it are known too; they are just held in people rather than files. The datasets that would matter are the ones recording the hand in motion, and one of them exists and is small and open and largely unused.

What is missing is not compute, or model capacity, or volume — this was never going to be fixed by more of what already exists. It is a different kind of record altogether, one nobody has made, of a body of knowledge whose entire transmission mechanism was designed around not needing one. Tarteel got 67 hours of Qur’anic recitation from twelve hundred volunteers in six months, which suggests the community can build the corpus when someone frames the task. For calligraphy, nobody has framed it.

Several things I could not settle, and would take a correction on. Whether anyone has run the stroke-generation playbook on Calliar and simply not published it. How much of DecoType’s model is documented well enough to rebuild from, now that the product is gone. Whether there is Arabic robotic calligraphy work in a language I do not read: the literature on brush robots is largely Chinese, and I searched in English and Arabic only. And how far the disputed figures go: sources do not even agree on how many dots high a Thuluth alif is, which is a strange place for a thousand-year-old specification to be vague.

What I am fairly confident of is the shape of it. The window is not open indefinitely: the chain that carries this knowledge runs through living people, and what they hold is the correction — the adjustment made to a hand in a room, which no photograph of a finished panel has ever contained. Every archive being built right now is an archive of dried ink.

The thing worth recording is the stroke.

Topics