API च्या किमतींमधील बदल क्वचितच लक्ष वेधून घेतात. यासाठी कोणताही स्टेटस-पेज अलर्ट, डेप्रिकेशन वॉर्निंग (deprecation warning) किंवा ईमेल येत नाही. प्रदाताच्या (provider) प्राइसिंग पेजवरील आकडेवारी फक्त बदलत जाते आणि पुढच्या वेळी जेव्हा तुमचे बॅच जॉब पूर्ण होतात, तेव्हा बिल वेगळे दिसते. Novita आणि StreamLake सोबत नेमके हेच घडले आहे. दोन्ही प्लॅटफॉर्मनी त्यांच्या LLM रेट कार्ड्समध्ये बदल केले आहेत, आणि जर तुम्ही यापैकी कोणत्याही सेवेवर इन्फरन्स वर्कलोड्स (inference workloads) चालवत असाल, तर तुमचे पुढचे काम सुरू करण्यापूर्वी तुम्हाला नवीन आकडेवारी तपासून पाहणे आवश्यक आहे.

बदलत्या किमतींचे शांत संकट

बहुतेक इंजिनिअरिंग टीम्स अपटाइम (uptime), लेटन्सी (latency) आणि टोकन अचूकता (token accuracy) अत्यंत बारकाईने मॉनिटर करतात. मात्र, प्रति हजार टोकनचा खर्च (cost per thousand tokens) सहसा ऑनबोर्डिंगच्या वेळी एकदाच पाहिला जातो आणि नंतर तो मागे पडतो. ही एक चूक आहे. हाय-व्हॉल्यूम ॲप्लिकेशन्समध्ये—जसे की कस्टमर सपोर्ट चॅटबॉट्स, डॉक्युमेंट समरायझेशन पाईपलाईन्स, कोड जनरेशन टूल्स—प्रति टोकनच्या किमतीत अगदी थोडासा बदल झाला तरी महिन्याच्या अखेरीस तो बजेटवर मोठा परिणाम करू शकतो.

फीचर डेप्रिकेशनच्या (feature deprecation) उलट, ज्यामुळे त्वरित कोडमध्ये बदल करावा लागतो, प्राइसिंग अपडेटमुळे तुमच्या इंटिग्रेशनवर कोणताही परिणाम होत नाही. तुमच्या विनंत्या (requests) अजूनही 200 स्टेटस कोड्स परत करतात. तुमचे JSON पेलोड्स अजूनही योग्य दिसतात. एकमेव फरक म्हणजे इन्व्हॉइस (invoice). जेव्हा फायनान्स टीमला या फरकाची जाणीव होते, तोपर्यंत तुम्ही कदाचित एका संपूर्ण स्प्रिंटचा इन्फरन्स बजेट खर्च केलेला असू शकतो. Novita आणि StreamLake या दोन्ही कंपन्यांनी अलीकडेच त्यांच्या प्राइसिंग स्ट्रक्चरमध्ये बदल केले आहेत, याचा अर्थ असा की त्यांच्या एंडपॉइंट्सवर (endpoints) काम करणारी कोणतीही ऑटोमेटेड पाईपलाईन, स्टेजिंग टेस्ट किंवा प्रोडक्शन वर्कलोड तुमच्या अपेक्षेपेक्षा जास्त किंवा कमी खर्चिक असू शकतो. अंदाज लावणे ही कोणतीही रणनीती नाही.

ताज्या अपडेट्सबद्दल आपल्याला काय माहिती आहे

Novita आणि StreamLake साठी प्रसिद्ध करण्यात आलेली रेट कार्ड्स बदलली आहेत. जरी नेमके बदल मॉडेल टियर आणि टोकन प्रकारानुसार वेगवेगळे असले, तरी मुख्य निष्कर्ष एकच आहे: इन्फरन्स खर्चाबाबत गेल्या महिन्यात तुम्ही केलेले अंदाज आता लागू पडणार नाहीत. Novita, जे GPU क्लाउड सेवांसोबत विविध लार्ज लँग्वेज मॉडेल (LLM) APIs प्रदान करते, त्याने मॉडेल ॲक्सेससाठी आकारले जाणारे शुल्क समायोजित केले आहे. StreamLake, जे क्लाउड आणि AI इन्फ्रास्ट्रक्चर प्रदाता म्हणून काम करते, त्याने देखील त्याच्या LLM प्राइसिंग शेड्युलमध्ये सुधारणा केली आहे.

कारण हे प्लॅटफॉर्म खर्च वेगवेगळ्या प्रकारे स्ट्रक्चर करतात—काही इनपुट आणि आउटपुट टोकन्स वेगळे करतात, काही त्यांचे बंडल करतात, तर काही लाँग-कॉन्टेक्स्ट विंडोज (long-context windows) किंवा हाय-थ्रूपुट एंडपॉइंट्ससाठी प्रीमियम आकारतात—त्यामुळे तुम्ही जुन्या अंदाजाचा वापर नवीन कामासाठी सुरक्षितपणे करू शकत नाही. एखादा वर्कफ्लो जो सोमवारी किफायतशीर होता, तो बुधवारी खर्चिक ठरू शकतो जर आउटपुट-टोकन मल्टिप्लायर बदलला असेल किंवा डिस्काउंट टियरमध्ये बदल झाला असेल. विशिष्ट रेटमधील बदलांचा तपशील मूळ डेव्हलपर रिपोर्टमध्ये दस्तऐवजीकृत केला आहे. त्या रिपोर्टला तुम्ही तुमचा मुख्य आधार मानावा, कोणत्याही थर्ड-पार्टी सारांशाला नाही.

LLM रेट कार्ड कसे वाचावे

जुन्या खर्चाची तुलना नवीन खर्चाशी करण्यापूर्वी, तुम्हाला नेमके काय पाहायचे आहे हे माहित असणे आवश्यक आहे. बहुतेक प्रदाते किमतींचे काही विशिष्ट भागांत विभाजन करतात आणि Novita आणि StreamLake अपवाद नाहीत.

पहिले म्हणजे, इनपुट टोकन्स आणि आउटपुट टोकन्स वेगळे करा. इनपुट म्हणजे तुम्ही मॉडेलला पाठवलेली माहिती; आउटपुट म्हणजे मॉडेलद्वारे तयार केलेली माहिती. अनेक प्रोडक्शन सिस्टममध्ये, आउटपुटचे प्रमाण इनपुटपेक्षा जास्त असते, विशेषतः चॅट समरायझेशन किंवा क्रिएटिव्ह रायटिंग टास्कमध्ये. जो प्रदाता इनपुट खर्च कमी करतो पण आउटपुट खर्च वाढवतो, तो प्रत्यक्षात तुमचा एकूण बिल वाढवू शकतो.

दुसरे म्हणजे, कॉन्टेक्स्ट-विंडो (context-window) प्राइसिंगवर लक्ष द्या. लाँग-कॉन्टेक्स्ट मॉडेल्स, जे एकाच वेळी हजारो किंवा लाखो टोकन्स हाताळतात, त्यांच्यावर कधीकधी प्रीमियम आकारला जातो जो रेषीय (linear) पद्धतीने वाढत नाही. जर तुमचे ॲप्लिकेशन संपूर्ण कोडबेस किंवा लांब कायदेशीर कागदपत्रे प्रॉम्प्ट म्हणून पाठवत असेल, तर लाँग-कॉन्टेक्स्ट टियरमधील प्रति-टोकन वाढ सामान्य वाढीपेक्षा जास्त परिणामकारक ठरू शकते.

तिसरे म्हणजे, थ्रूपुट (throughput) आणि कन्करन्सी (concurrency) नियमांकडे लक्ष द्या. काही रेट कार्ड्स बॅच किंवा ऑफलाइन इन्फरन्ससाठी कमी किमती देतात, परंतु रिअल-टाइम स्ट्रीमिंगसाठी जास्त शुल्क आकारतात. जर तुमचे युजर-फेसिंग ॲप्लिकेशन कमी लेटन्सी प्रतिसादांवर अवलंबून असेल, तर टोकन व्हॉल्यूम कितीही असला तरी तुम्हाला प्रीमियम टियरमध्ये राहावे लागू शकते.

शेवटी, लपलेले अतिरिक्त खर्च (auxiliary costs) तपासा. रिट्रिव्हल-ऑगमेंटेड जनरेशन (RAG) पाईपलाईन्स अनेकदा LLM पर्यंत पोहोचण्यापूर्वी एम्बेडिंग एंडपॉइंट्स, वेक्टर स्टोअर्स आणि रँकिंग APIs चा वापर करतात. जरी Novita आणि StreamLake ने त्यांच्या LLM प्राइसिंगमध्ये अपडेट केले असले, तरी त्याच इन्व्हॉइसमधील संबंधित सेवांच्या किमती देखील बदलल्या असू शकतात. फक्त प्रति-मिलियन-टोकन रेटची हेडलाईन न वाचता संपूर्ण पेज वाचा.

तुमच्या पुढच्या डिप्लॉयमेंटपूर्वी आकडेवारी तपासा

Once you have the fresh rate card, do not estimate. Measure. Pull your last seven to thirty days of request logs and calculate what that identical workload would cost under the new structure. If you are using a centralized logging tool or an observability dashboard, filter by the provider endpoint and export token counts. Most APIs return usage metadata in the response payload, so you can script this in a few lines of Python.

Start with a representative sample. Pick your busiest day from the previous billing cycle. Multiply the input tokens by the new input rate and the output tokens by the new output rate. Add any context-window or throughput surcharges that apply to your model tier. Compare that synthetic bill against what you actually paid. If the delta crosses your tolerance threshold—say, ten or twenty percent—you have a decision to make.

That decision does not always mean migrating providers. Sometimes it means switching model tiers within the same platform, trimming prompt length, enabling response caching, or throttling non-critical batch jobs to off-peak hours. The point is to make that decision with data rather than discovering the change on the next invoice.

You should also set hard spend caps or budget alerts if the platform supports them. Many API dashboards allow you to configure notification thresholds at the project or key level. Place them conservatively. If Novita or StreamLake push another rate change in the future, you want a financial circuit breaker, not a surprise four-figure overage.

The bigger picture: infrastructure costs are never static

These updates from Novita and StreamLake are reminders that the foundation-model market is still settling. Pricing is not an accident; it reflects compute availability, licensing deals, and competitive positioning. A provider might cut rates to attract volume, then raise them once a user base is locked in. Alternatively, a provider might raise rates to cover the cost of newer, more capable models while grandfathering older ones. Either way, relying on a single provider’s rate card as a constant is poor operational hygiene.

Teams who treat inference as a commodity layer already run multi-provider setups. They route simple queries to the cheapest endpoint that meets a quality bar and reserve expensive models for hard tasks. That architecture requires more upfront wiring, but it insulates you from exactly this kind of quiet price shift. Even if you are not ready to deploy a full routing layer, keeping a secondary provider warm and benchmarked gives you leverage when the primary one moves its prices.

Where to find the exact figures

The granular breakdown of what changed—model by model, token type by token type—is available in the original report. You can read the full details at the source link that tracked these updates. For ongoing discussions about infrastructure pricing, model releases, and cost optimization tactics, the GyaanSetu learning community is active on Telegram.

The takeaway

Do not let a pricing update become a post-mortem. Before you queue up your next training run, batch inference job, or production deployment against Novita or StreamLake, open their current pricing pages and rerun your last week’s numbers against the new rates. If the math still works, proceed with confidence. If it does not, you have the data to renegotiate your pipeline before the meter starts running again. Your future invoice will thank you.