API-এর মূল্যের পরিবর্তন খুব কমই শোরগোল ফেলে। কোনো স্ট্যাটাস-পেজ অ্যালার্ট, কোনো ডিপ্রিকেশন ওয়ার্নিং (deprecation warning), এমনকি সাধারণত কোনো ইমেল ব্ল্যাস্টও পাঠানো হয় না। প্রোভাইডারের প্রাইসিং পেজের সংখ্যাগুলো কেবল পরিবর্তিত হয়, এবং পরের বার যখন আপনার ব্যাচ জব শেষ হয়, তখন বিলটি ভিন্ন দেখায়। Novita এবং StreamLake-এর ক্ষেত্রে ঠিক এটাই ঘটেছে। উভয় প্ল্যাটফর্মই তাদের LLM rate card আপডেট করেছে, এবং আপনি যদি এই পরিষেবাগুলোর কোনোটিতে ইনফারেন্স ওয়ার্কলোড (inference workloads) চালান, তবে আপনার পরবর্তী টাস্ক শুরু করার আগে নতুন সংখ্যাগুলো দেখে নেওয়া প্রয়োজন।

পরিবর্তনশীল মূল্যের লক্ষ্যমাত্রার নীরব হুমকি

বেশিরভাগ ইঞ্জিনিয়ারিং টিম অত্যন্ত নিষ্ঠার সাথে আপটাইম (uptime), ল্যাটেন্সি (latency) এবং টোকেন নির্ভুলতা (token accuracy) পর্যবেক্ষণ করে। তবে, প্রতি হাজার টোকেনের খরচ সাধারণত অনবোর্ডিংয়ের সময় একবার দেখা হয় এবং তারপর তা অবহেলিত থেকে যায়। এটি একটি ভুল। উচ্চ-ভলিউম অ্যাপ্লিকেশন—যেমন কাস্টমার সাপোর্ট চ্যাটবট, ডকুমেন্ট সামারাইজেশন পাইপলাইন, কোড জেনারেশন টুল—এগুলোতে প্রতি টোকেনে সামান্যতম পরিবর্তনও মাসের শেষে বাজেটে বড় ধরনের চাপ সৃষ্টি করতে পারে।

ফিচার ডিপ্রিকেশন (feature deprecation)-এর মতো নয়, যা তাৎক্ষণিক কোড পরিবর্তনের বাধ্যবাধকতা তৈরি করে, একটি প্রাইসিং আপডেট আপনার ইন্টিগ্রেশনকে অপরিবর্তিত রাখে। আপনার রিকোয়েস্টগুলো এখনও 200 status code রিটার্ন করে। আপনার JSON payloads এখনও সঠিক দেখায়। একমাত্র পার্থক্য হলো ইনভয়েস বা বিল। ফিন্যান্স বিভাগ যখন অসংগতিটি শনাক্ত করে, ততক্ষণে আপনি হয়তো একটি পুরো স্প্রিন্টের (sprint) ইনফারেন্স বাজেট শেষ করে ফেলেছেন। Novita এবং StreamLake উভয়ই সম্প্রতি তাদের প্রাইসিং স্ট্রাকচার পরিবর্তন করেছে, যার মানে হলো তাদের এন্ডপয়েন্টে হিট করা যেকোনো স্বয়ংক্রিয় পাইপলাইন, স্টেজিং টেস্ট বা প্রোডাকশন ওয়ার্কলোড আপনার প্রত্যাশার চেয়ে বেশি বা কম খরচ করতে পারে। অনুমানের ওপর নির্ভর করা কোনো কৌশল নয়।

সাম্প্রতিক আপডেট সম্পর্কে আমরা যা জানি

Novita এবং StreamLake-এর প্রকাশিত rate card উভয়ই পরিবর্তিত হয়েছে। যদিও মডেল টিয়ার এবং টোকেন টাইপের ওপর ভিত্তি করে সঠিক পার্থক্য ভিন্ন ভিন্ন হয়, তবে মূল বিষয়টি একই: ইনফারেন্স খরচের বিষয়ে গত মাসে আপনার যে ধারণা ছিল, তা আর নাও থাকতে পারে। Novita, যা GPU ক্লাউড সার্ভিসের পাশাপাশি বিভিন্ন ধরনের Large Language Model API অফার করে, তারা মডেল অ্যাক্সেসের চার্জ করার পদ্ধতিতে পরিবর্তন এনেছে। StreamLake, যা একটি বিস্তৃত ক্লাউড এবং AI ইনফ্রাস্ট্রাকচার প্রোভাইডার হিসেবে কাজ করে, তারাও একইভাবে তাদের LLM প্রাইসিং শিডিউল সংশোধন করেছে।

যেহেতু এই প্ল্যাটফর্মগুলো খরচ ভিন্ন ভিন্নভাবে সাজায়—কেউ ইনপুট এবং আউটপুট টোকেন আলাদা করে, কেউ সেগুলোকে বান্ডেল করে, আবার কেউ লং-কনটেক্সট উইন্ডো বা হাই-থ্রুপুট এন্ডপয়েন্টের জন্য অতিরিক্ত চার্জ যোগ করে—তাই আপনি একটি পুরনো হিসাবকে নিরাপদে নতুন কোনো কাজের ওপর প্রয়োগ করতে পারবেন না। সোমবার যে ওয়ার্কফ্লো সাশ্রয়ী ছিল, বুধবার তা ব্যয়বহুল হয়ে উঠতে পারে যদি আউটপুট-টোকেন মাল্টিপ্লায়ার পরিবর্তিত হয় বা কোনো ডিসকাউন্ট টিয়ার পুনর্গঠিত হয়। নির্দিষ্ট রেট পরিবর্তনের বিস্তারিত বিবরণ মূল ডেভেলপার রিপোর্টে নথিভুক্ত করা আছে। সেই রিপোর্টটিকে আপনার মূল সত্য (ground truth) হিসেবে বিবেচনা করুন, কোনো থার্ড-পার্টি সারাংশ হিসেবে নয়।

কীভাবে একটি LLM rate card পড়তে হয়

পুরানো খরচের সাথে নতুন খরচের তুলনা করার আগে, আপনি আসলে কী দেখছেন তা জানা প্রয়োজন। বেশিরভাগ প্রোভাইডার প্রাইসিংকে কয়েকটি স্বতন্ত্র উপাদানে ভাগ করে, এবং Novita ও StreamLake এর ব্যতিক্রম নয়।

প্রথমত, ইনপুট টোকেন থেকে আউটপুট টোকেন আলাদা করুন। ইনপুট হলো যা আপনি মডেলে পাঠান; আউটপুট হলো যা মডেল তৈরি করে। অনেক প্রোডাকশন সিস্টেমে, আউটপুট ভলিউম ইনপুট ভলিউমের চেয়ে বেশি হয়, বিশেষ করে চ্যাট সামারাইজেশন বা ক্রিয়েটিভ রাইটিং টাস্কের ক্ষেত্রে। একজন প্রোভাইডার যদি ইনপুট খরচ কমায় কিন্তু আউটপুট খরচ বাড়ায়, তবে তা প্রকৃতপক্ষে আপনার মোট বিল বাড়িয়ে দিতে পারে।

দ্বিতীয়ত, কনটেক্সট-উইন্ডো প্রাইসিং (context-window pricing)-এর দিকে নজর দিন। লং-কনটেক্সট মডেলগুলো, যা একটি সিঙ্গেল পাসে দশ বা লক্ষাধিক টোকেন হ্যান্ডেল করে, কখনও কখনও অতিরিক্ত প্রিমিয়াম চার্জ করে যা রৈখিকভাবে (linearly) বাড়ে না। যদি আপনার অ্যাপ্লিকেশন প্রম্পট হিসেবে সম্পূর্ণ কোডবেস বা দীর্ঘ আইনি নথি পাঠায়, তবে লং-কনটেক্সট টিয়ারে প্রতি-টোকেন সামান্য বৃদ্ধিও সাধারণ বৃদ্ধির চেয়ে বেশি প্রভাব ফেলে।

তৃতীয়ত, থ্রুপুট (throughput) এবং কনকারেন্সি (concurrency) সংক্রান্ত নিয়মগুলো দেখুন। কিছু rate card ব্যাচড (batched) বা অফলাইন ইনফারেন্সের জন্য কম দাম অফার করে কিন্তু রিয়েল-টাইম স্ট্রিমিংয়ের জন্য বেশি চার্জ করে। যদি আপনার ইউজার-ফেসিং অ্যাপ্লিকেশন লো-ল্যাটেন্সি রেসপন্সের ওপর নির্ভর করে, তবে টোকেন ভলিউম যাই হোক না কেন, আপনি একটি প্রিমিয়াম টিয়ারে আটকে থাকতে পারেন।

সবশেষে, লুকানো সহায়ক খরচগুলো (auxiliary costs) পরীক্ষা করুন। রিট্রিভাল-অগমেন্টেড জেনারেশন (RAG) পাইপলাইনগুলো LLM-এ পৌঁছানোর আগে প্রায়শই এমবেডিং এন্ডপয়েন্ট (embedding endpoints), ভেক্টর স্টোর (vector stores) এবং র র‍্যাঙ্কিং (reranking) API ব্যবহার করে। যদিও Novita এবং StreamLake তাদের LLM প্রাইসিং আপডেট করেছে, তবে একই ইনভয়েসে থাকা পার্শ্ববর্তী পরিষেবাগুলোর দামও পরিবর্তিত হতে পারে। শুধুমাত্র প্রতি-মিলিয়ন-টোকেন রেট বা শিরোনামটি না দেখে পুরো পৃষ্ঠাটি পড়ুন।

আপনার পরবর্তী ডিপ্লয়মেন্টের আগে হিসাব করে নিন

Once you have the fresh rate card, do not estimate. Measure. Pull your last seven to thirty days of request logs and calculate what that identical workload would cost under the new structure. If you are using a centralized logging tool or an observability dashboard, filter by the provider endpoint and export token counts. Most APIs return usage metadata in the response payload, so you can script this in a few lines of Python.

Start with a representative sample. Pick your busiest day from the previous billing cycle. Multiply the input tokens by the new input rate and the output tokens by the new output rate. Add any context-window or throughput surcharges that apply to your model tier. Compare that synthetic bill against what you actually paid. If the delta crosses your tolerance threshold—say, ten or twenty percent—you have a decision to make.

That decision does not always mean migrating providers. Sometimes it means switching model tiers within the same platform, trimming prompt length, enabling response caching, or throttling non-critical batch jobs to off-peak hours. The point is to make that decision with data rather than discovering the change on the next invoice.

You should also set hard spend caps or budget alerts if the platform supports them. Many API dashboards allow you to configure notification thresholds at the project or key level. Place them conservatively. If Novita or StreamLake push another rate change in the future, you want a financial circuit breaker, not a surprise four-figure overage.

The bigger picture: infrastructure costs are never static

These updates from Novita and StreamLake are reminders that the foundation-model market is still settling. Pricing is not an accident; it reflects compute availability, licensing deals, and competitive positioning. A provider might cut rates to attract volume, then raise them once a user base is locked in. Alternatively, a provider might raise rates to cover the cost of newer, more capable models while grandfathering older ones. Either way, relying on a single provider’s rate card as a constant is poor operational hygiene.

Teams who treat inference as a commodity layer already run multi-provider setups. They route simple queries to the cheapest endpoint that meets a quality bar and reserve expensive models for hard tasks. That architecture requires more upfront wiring, but it insulates you from exactly this kind of quiet price shift. Even if you are not ready to deploy a full routing layer, keeping a secondary provider warm and benchmarked gives you leverage when the primary one moves its prices.

Where to find the exact figures

The granular breakdown of what changed—model by model, token type by token type—is available in the original report. You can read the full details at the source link that tracked these updates. For ongoing discussions about infrastructure pricing, model releases, and cost optimization tactics, the GyaanSetu learning community is active on Telegram.

The takeaway

Do not let a pricing update become a post-mortem. Before you queue up your next training run, batch inference job, or production deployment against Novita or StreamLake, open their current pricing pages and rerun your last week’s numbers against the new rates. If the math still works, proceed with confidence. If it does not, you have the data to renegotiate your pipeline before the meter starts running again. Your future invoice will thank you.