Where the AI bill actually goesأين تذهب فاتورة الذكاء الاصطناعي فعليًا

Token prices fell by roughly 80% in a year and enterprise AI bills went up anyway. Four things we found in our own numbers, running around 1.3 million chat turns a month — starting with the discovery that almost none of the cost is the part doing the thinking.انخفضت أسعار التوكِن نحو ٨٠٪ خلال عام، ومع ذلك ارتفعت فواتير الذكاء الاصطناعي في المؤسسات. أربع ملاحظات من أرقامنا نحن، مع تشغيل ما يقارب ١٫٣ مليون محادثة شهريًا — أولها أن الجزء «الذكي» ليس هو التكلفة تقريبًا.

Two facts that do not want to sit together. Token prices dropped by something close to 80% between early 2025 and early 2026. Over the same period enterprise spend on inference went up, and survey after survey finds most AI programmes overshooting their cost estimate by a third to a half.حقيقتان لا تستقران معًا بسهولة. انخفضت أسعار التوكِن بما يقارب ٨٠٪ بين مطلع ٢٠٢٥ ومطلع ٢٠٢٦، وفي المدة نفسها ارتفع إنفاق المؤسسات على الاستدلال، وتتكرر الدراسات التي تجد أن أغلب برامج الذكاء الاصطناعي تتجاوز تقديرها للتكلفة بالثلث إلى النصف.

The resolution is unglamorous: consumption grows faster than unit prices fall. But it leaves an operator with a question no pricing page answers — inside a contact centre that is actually running, which part of the bill is large?التفسير غير مثير: الاستهلاك ينمو أسرع من انخفاض سعر الوحدة. لكنه يترك سؤالًا لا تجيب عنه صفحات الأسعار — داخل مركز اتصال يعمل فعلًا، أي جزء من الفاتورة هو الكبير؟

We went looking in ours. Four findings, in the order they surprised us.بحثنا في أرقامنا. أربع ملاحظات، بترتيب ما فاجأنا منها.

Almost none of it is the outputالإخراج لا يكاد يمثّل شيئًا

We broke a month of platform token usage into three buckets. The split was not close.قسّمنا استهلاك شهر كامل من التوكِن على ثلاث فئات، وجاءت النتيجة غير متقاربة إطلاقًا.

Share of token spend, one monthتوزيع إنفاق التوكِن خلال شهر
Cached inputمدخلات مخزّنة86%
Fresh inputمدخلات جديدة13.5%
Outputمخرجات0.5%

Nearly every cost-cutting instinct in the industry points at the bottom row.ومعظم غرائز خفض التكلفة في السوق تستهدف الصف الأخير.

Shortening answers. Capping response length. Telling the model to be concise. All of it operates on half a percent of the bill. The system prompt, the tool schemas, the retrieved context and the conversation history are the bill.تقصير الإجابات، وتحديد سقف لطول الرد، ومطالبة النموذج بالإيجاز — كل ذلك يعمل على نصف بالمئة من الفاتورة. أما التعليمات الأساسية ومخططات الأدوات والسياق المسترجَع وتاريخ المحادثة، فهي الفاتورة نفسها.

Two things follow. The highest-leverage work is confirming your provider’s prompt caching is genuinely being hit, because that 86% is only cheap while the cache holds — a change that quietly busts it does not look like an incident, it looks like a larger invoice. And the thing to review is not how the model writes. It is what you send it on every single turn.ينتج عن ذلك أمران. أعلى الأعمال أثرًا هو التأكد من أن تخزين التعليمات المؤقت يعمل فعلًا، لأن تلك الـ٨٦٪ تبقى رخيصة ما دام التخزين صامدًا — وأي تغيير يكسره بهدوء لا يبدو كعُطل، بل يبدو كفاتورة أكبر. والأمر الثاني أن ما يستحق المراجعة ليس كيف يكتب النموذج، بل ما ترسله إليه في كل دورة.

A large share of turns are not questionsجزء كبير من الرسائل ليس أسئلة

The second finding needed no instrumentation to understand. A lot of what arrives in a conversation is not a request for anything: greetings, thanks, “ok”, a thumbs-up, the acknowledgement at the end of a thread that is already resolved.الملاحظة الثانية لا تحتاج أدوات قياس لفهمها. كثير مما يصل في المحادثة ليس طلبًا لشيء: تحيات، وشكر، و«تمام»، وعلامة إعجاب، وإقرار ختامي في محادثة انتهت أصلًا.

Every one of those was going through the full pipeline — retrieval across the knowledge base, the agent loop, then a frontier model composing a courteous reply. We were paying the most expensive component in the stack to recognise the word thanks.كل ذلك كان يمرّ عبر المسار الكامل — بحث في قاعدة المعرفة، ثم حلقة الوكيل، ثم نموذج متقدّم يصوغ ردًا مهذبًا. كنا ندفع لأغلى مكوّن في المنظومة كي يتعرّف على كلمة شكرًا.

There is now a classification step in front of it. A small open-weights model decides whether the turn is an acknowledgement or a real question. Acknowledgements return a signal instead of an answer, before retrieval or generation runs at all. Real questions pass through untouched.صارت هناك الآن خطوة تصنيف قبله. نموذج صغير مفتوح الأوزان يحدّد إن كانت الرسالة إقرارًا أم سؤالًا حقيقيًا. الإقرارات تُعيد إشارة بدل إجابة، قبل تشغيل البحث أو التوليد أصلًا. أما الأسئلة الحقيقية فتمرّ كما هي.

Three details matter more than the idea itself:وثلاث تفاصيل هنا أهم من الفكرة نفسها:

  • It fails open. Any error or timeout and the customer gets a normal answer. A cost optimisation that can break a conversation is not a cost optimisation.يفشل مفتوحًا. أي خطأ أو انتهاء مهلة يعني أن العميل يحصل على إجابة عادية. تحسين التكلفة الذي قد يكسر محادثة ليس تحسينًا للتكلفة.
  • It costs latency — 350 to 500 milliseconds on the turns it classifies. That is a real trade, and worth stating plainly rather than burying.له كلفة في الزمن — من ٣٥٠ إلى ٥٠٠ جزء من الألف من الثانية على الرسائل التي يصنّفها. مقايضة حقيقية تستحق أن تُقال صراحة لا أن تُخفى.
  • It was validated in both languages before it went anywhere near production. An acknowledgement classifier that only really understands English will quietly start swallowing questions in Arabic, and nothing in your dashboards will say so.جرى التحقق منه باللغتين قبل اقترابه من الإنتاج. مصنّف إقرارات يفهم الإنجليزية وحدها سيبدأ بابتلاع الأسئلة العربية بهدوء، ولن تخبرك بذلك أي لوحة مؤشرات.

“Cheap” is not one price band«الرخيص» ليس فئة سعرية واحدة

Choosing the model for that classifier took longer than building it, and produced the most portable lesson here.اختيار نموذج هذا المصنّف استغرق وقتًا أطول من بنائه، وأنتج أكثر الدروس قابلية للنقل.

Between three models that all qualify as small and cheap, the annual cost of running the same classification at our volume differed by a factor of five. Not five percent. Five times. Every one of them reads as inexpensive on a pricing page.بين ثلاثة نماذج تُوصف كلها بالصغيرة والرخيصة، اختلفت التكلفة السنوية لتشغيل التصنيف نفسه عند حجمنا بمقدار خمسة أضعاف. ليس خمسة بالمئة. خمسة أضعاف. وكلها تبدو زهيدة على أي صفحة أسعار.

At low volume that gap is invisible and not worth an afternoon. At a million-plus turns a month it is a line item. The only way to see it is to multiply list price by your own measured token shape — ours averages about 431 input tokens and 2 output tokens per classification, a profile so lopsided that input pricing decides everything and output pricing is close to irrelevant.عند الأحجام الصغيرة يكون هذا الفارق غير مرئي ولا يستحق وقتًا. وعند أكثر من مليون رسالة شهريًا يصبح بندًا في الميزانية. والسبيل الوحيد لرؤيته هو ضرب السعر المعلن في شكل التوكِن الخاص بك — متوسطنا نحو ٤٣١ توكِن مدخلات ومقابل توكِنين للمخرجات، وهو توزيع مائل إلى حد يجعل سعر المدخلات هو الحاسم وسعر المخرجات شبه بلا أثر.

Your shape will be different. That is precisely the point: the ranking of models by cost depends on your traffic, not on theirs.شكل بياناتك سيكون مختلفًا. وهذا هو المقصود تحديدًا: ترتيب النماذج من حيث التكلفة يعتمد على حركتك أنت، لا على أمثلتهم.

Cheaper and better stopped being oppositesالأرخص والأفضل لم يعودا نقيضين

The old reflex was that cutting model cost meant accepting worse answers. That is no longer reliably true. Our most recent migration moved the platform onto a newer frontier model that took roughly 86% off the cost of a typical turn while winning every shared benchmark we compared the two on, with a context window more than twice as large.كانت القاعدة القديمة أن خفض تكلفة النموذج يعني القبول بإجابات أسوأ. لم يعد ذلك صحيحًا على الدوام. آخر انتقال لنا نقل المنصة إلى نموذج متقدّم أحدث خفّض نحو ٨٦٪ من تكلفة الرسالة النموذجية، وتفوّق في كل الاختبارات المرجعية المشتركة التي قارنّا بها، بنافذة سياق تزيد على الضعف.

The interesting decision was not the migration. It was the scope. We moved 492 prompts and deliberately left several categories where they were: voice and IVR, the query-rewriting step inside retrieval, media recognition, and the live chat of our largest customer — where the cost of a subtle regression is worth more than the saving. Summarisation and quality scoring for that customer moved. The conversations their customers actually see did not.لكن القرار المثير للاهتمام لم يكن الانتقال، بل نطاقه. نقلنا ٤٩٢ تعليمة، وتركنا عمدًا فئات في مكانها: الصوت ونظام الرد الآلي، وخطوة إعادة صياغة الاستعلام داخل البحث، والتعرّف على الوسائط، والمحادثة المباشرة لأكبر عملائنا — حيث تكلفة أي تراجع طفيف تفوق قيمة التوفير. أما التلخيص وتقييم الجودة لذلك العميل فقد انتقلا. والمحادثات التي يراها عملاؤه فعلًا لم تنتقل.

Move the parts where a regression is recoverable first. Let them run. Move the rest on evidence rather than on a benchmark table.انقل أولًا الأجزاء التي يمكن التراجع عن أخطائها. دعها تعمل. وانقل الباقي بناءً على دليل لا على جدول اختبارات.

The failure mode nobody warns you aboutنمط الفشل الذي لا يحذّرك منه أحد

The new model needed four configuration fields set together: model, provider, reasoning effort, temperature. Set three of them and the prompt does not fail. It works — and becomes quietly slower and more expensive.احتاج النموذج الجديد إلى أربعة حقول إعداد تُضبط معًا: النموذج، والمزوّد، ومستوى الاستدلال، ودرجة العشوائية. اضبط ثلاثة منها فقط ولن تفشل التعليمة. ستعمل — وتصير أبطأ وأغلى بهدوء.

That is the shape of almost every cost incident we have seen. Nothing goes down. No error rate moves. No customer complains. The bill arrives three weeks later.هذا هو شكل معظم حوادث التكلفة التي رأيناها. لا شيء يتوقف. لا معدل خطأ يتحرك. لا عميل يشتكي. وتصل الفاتورة بعد ثلاثة أسابيع.

So if there is one operational habit to take from this: alarm on cost per conversation, not on total spend. Total spend moves with traffic and hides everything inside it. Cost per conversation catches a misconfiguration in days instead of at the end of the quarter.فإن كان ثمة عادة تشغيلية واحدة تؤخذ من هذا: اضبط تنبيهًا على تكلفة المحادثة الواحدة، لا على الإنفاق الإجمالي. الإنفاق الإجمالي يتحرك مع حجم الحركة ويخفي كل شيء بداخله. أما تكلفة المحادثة فتكشف أي خطأ إعداد خلال أيام بدل نهاية الربع.

And the model is only one of the metersوالنموذج ليس سوى أحد العدّادين

Worth saying, because it surprises people in this region more than any of the above: there are two bills running, and they behave nothing alike. On WhatsApp, Meta charges per delivered message at a rate set by the recipient’s market and the message category, with graduated volume tiers — and for most of our customers that bill is the larger of the two. We built a calculator for it.يستحق الذكر لأنه يفاجئ الناس في منطقتنا أكثر من كل ما سبق: هناك فاتورتان تعملان، وسلوكهما مختلف تمامًا. على واتساب تحتسب ميتا التكلفة لكل رسالة تُسلَّم، بسعر يحدده سوق المستلم وفئة الرسالة، مع شرائح حجم متدرجة — ولدى معظم عملائنا تكون تلك الفاتورة هي الأكبر. وقد بنينا حاسبة لها.

None of this is exotic. Measure your token shape before optimising anything. Don’t send expensive models work a cheap one can refuse. Price the cheap options against your own traffic, not the vendor’s example. Migrate the recoverable parts first. And put an alarm on cost per conversation, because in this domain the expensive mistakes are the quiet ones.لا شيء من هذا استثنائي. قِس شكل التوكِن لديك قبل أن تحسّن أي شيء. ولا ترسل إلى نماذج غالية عملًا يستطيع نموذج رخيص رفضه. وسعّر الخيارات الرخيصة مقابل حركتك أنت لا مقابل مثال المزوّد. وانقل الأجزاء القابلة للتراجع أولًا. واضبط تنبيهًا على تكلفة المحادثة، لأن الأخطاء المكلفة في هذا المجال هي الأخطاء الصامتة.

See it run on your own traffic.شاهدها تعمل على حركتك أنت.

Thirty minutes with a CX engineer. No slideware.ثلاثون دقيقة مع مهندس تجربة عملاء. بلا شرائح عرض.