Skip to main content
Blog

Arabic NLP: Challenges and Practical Solutions

28/03/1448 AH

09/09/2026

Arabic breaks many NLP pipelines that work fine in English. The core issue is diglossia plus morphology plus orthographic variation, all at once. Customers in Dammam write Saudi dialect on WhatsApp, your docs are in Modern Standard Arabic, and your model was trained mostly on news MSA. Retrieval fails, generation leaks into Egyptian dialect, and search misses.

This is the full playbook for production Arabic NLP in Saudi apps in 2026, from normalization to models to RAG to evaluation.

1. Why Arabic is hard

Diglossia means MSA for formal writing versus dialects for speech and chat. Saudi Najdi uses وش تبي, Hijazi uses ايش تبغى, Egyptian uses عايز ايه for the same intent. A bot trained only on MSA misunderstands all three.

Morphology means one root generates hundreds of surface forms through prefixes, suffixes, and infixes. The root ك-ت-ب produces كتاب, كاتب, مكتوب, يكتبون, سنكتبها, and more. English tokenizers split these badly, inflating token counts 30 to 60 percent versus English and fragmenting meaning.

Orthography adds hamza variants of alef, alef maqsura versus ya, ta marbuta versus ha, tatweel for justification, diacritics that change meaning, and code-switching with English product names and Arabizi like 3shan for عشان. Two strings that look identical to users can be different Unicode sequences.

2. Normalization that preserves meaning

Normalize for retrieval but store original for display. Steps are unify alef forms to bare alef except when hamza changes meaning you care about, map alef maqsura to ya for search, strip tatweel, normalize Eastern and Western numerals, remove or standardize diacritics depending on domain, and clean ZWJ and control characters.

Keep both columns in your database, original_text for UI and normalized_text for search and embeddings. For names and VAT records, be conservative. Over-normalization merges distinct entities.

Build a Saudi dialect dictionary for your domain. Map وش, ايش, وشلون, ابغى, ابي to intents in your classifier training data. Collect 500 real support messages and label them. This small dataset beats any generic benchmark.

3. Models that work in 2026

For understanding Saudi dialect, Saudi ALLaM, Jais, and AceGPT outperform generic multilingual models on our tests. They handle code-switching and Gulf vocabulary better. For high-stakes reasoning in Arabic plus English, GPT-4 class models still lead but cost more and send data off-premise.

For embeddings, multilingual-e5-large with a reranker gives best Arabic retrieval in our RAG tests. Test with bge-reranker or Cohere rerank for Arabic. Chunk at 400 to 600 tokens respecting headings, keep title and section metadata, and test retrieval with dialect queries, not MSA translations of English benchmarks.

For speech, Whisper large with Arabic fine-tuning plus a Saudi dialect adapter handles WhatsApp voice notes. Post-process numbers and prices, since يا اخي ابغى ثلاث قطع بخمسين often mis-transcribes numerals.

4. Tokenization and cost control

Arabic costs more tokens. Measure fertility on your corpus. If your tokenizer uses 2x tokens for Arabic versus English, your RAG context and bills double. Mitigate by retrieving fewer but better chunks via reranking, caching embeddings, and summarizing long threads before LLM calls.

For generation, constrain output length in Arabic words, not tokens, and ask for concise replies. Provide few-shots in Saudi tone to avoid verbose MSA essays when users want a two-line answer.

5. RAG for Arabic docs

Ingest help articles as markdown with headings intact. Split by heading then by size, overlapping 10 percent. Embed normalized text but cite original. Retrieve top 8 with hybrid vector plus full-text, rerank to 3, inject with citations, and refuse when scores are low.

Prompt in Arabic with explicit dialect rule. Example system prompt states you are a Saudi support agent, answer in Saudi-friendly MSA, no Egyptian words like عايز or كده, prices in SAR, no invented policies. Provide three examples of good answers from your team.

Evaluate with 100 real questions from tickets covering dialect, prices, returns, and Fatoora. Score faithfulness, dialect leakage, and citation accuracy weekly. Log bad answers with user correction for retraining.

6. RTL and UX

Arabic UI needs dir rtl, logical CSS properties, mirrored icons, and Arabic numerals option. Chat bubbles should align right for user and left for agent in RTL. Dates should support Hijri alongside Gregorian for invoices and appointments. Test with VoiceOver in Arabic.

Bottom line is collect Saudi data, normalize carefully, choose dialect-aware models, and evaluate weekly. Generic Arabic demos impress, but Saudi production needs dialect work.

Innovative Solutions, Exceptional Results
Sikka Software © 2026
v2.18.3
madavisamastercardapple_paypaypalbank_transfer