Arabic prompts fail when they are vague about dialect, tone, and constraints. A prompt that says write about our services produces MSA essays with Egyptian words, English prices, and no call to action. A prompt that specifies Saudi audience in Dammam, formal-friendly tone, 150 words, SAR prices, and two variants produces usable drafts.
This is the full system for Arabic prompts that survive production, from templates to few-shots to evaluation.
1. Anatomy of a working Arabic prompt
Role must name Saudi copywriter, not generic writer. Audience must name SME owners in Dammam or Riyadh with their pain, not general readers. Tone must say formal-friendly Saudi MSA and explicitly ban Egyptian and Levantine words like عايز, كده, بدي. List banned words. Models respect explicit bans better than vague use Saudi dialect.
Task must state output shape. Title 60 characters plus 150-word body plus 3-bullet offer plus WhatsApp CTA. Constraints must state SAR VAT-inclusive, Arabic numerals option, no English except product names like Moyasar, and no invented discounts.
Example skeleton is role, audience, tone with bans, task with lengths, constraints with prices and CTA, then three examples, then request for two variants labeled A and B.
2. Few-shots that fix dialect
Provide three examples of input brief to ideal output. One for services, one for ecommerce product, one for Fatoora notice. Each shows Saudi phrasing like نخدمكم في الدمام, نوصل لجدة خلال يومين, السعر شامل الضريبة. Include one bad example marked do not do this with Egyptian leakage, and explain why. Models learn bans faster from contrast.
Keep examples versioned with dates. When the team approves a great Instagram caption, add it as a fourth shot and retire the weakest. Prompt libraries rot without gardening.
3. Constraints that prevent hallucination
Prices must cite a price list version. Availability must cite stock feed. Policies must cite doc slug. Instruct the model to write السعر حسب قائمة 2026-09 if unsure is forbidden. Better to output يحتاج تأكيد بشري than invent 299 SAR.
Length control works better in Arabic words than tokens. Say 150 كلمة عربية, not 200 tokens, since Arabic fertility varies. Ask for short sentences under 20 words for mobile readability.
4. Evaluation for Arabic quality
Score four axes on 20 samples weekly. Dialect leakage via banned-word list plus human spot check. Grammar via native review, not automated metrics that miss Saudi nuance. Factuality via price and policy match to source. CTA presence via regex for WhatsApp number and verb.
Track win rate of variant A versus B in real posts. Keep prompts in git with scores and dates. Small constraint tweaks like add city name often fix large gaps without model switching.
Test with real briefs from Dammam and Riyadh clients, not MSA benchmarks. A prompt that scores 90 on generic MSA but fails Saudi WhatsApp tone is useless.
5. Operations and versioning
Store prompts with version, owner, last eval date, and approved examples. Review monthly. Retire prompts that produce repetitive openers like في عالم اليوم. Ban stock openers explicitly and provide five approved openers.
Cost tip is cache system prompts and few-shots with prompt caching, stream UI, and use small models for first drafts then large for polish. This cuts cost 40 percent while keeping Saudi tone.
Bottom line is specificity wins. Name the audience, ban the dialects, constrain prices, show three Saudi examples, and evaluate weekly. Generic Arabic prompts produce generic results.





