גוגל ואנתרופיק מזהירות: סוכני ReAct עלולים לרוקן תקציבי ענן
לולאת ה-ReAct מאפשרת למודל לבחור את הצעדים שלו, אך כל מחזור של מחשבה-פעולה-תצפית מוסיף עלות הסקה נוספת, מה שמגדיל את השיהוי ומעלה את הסיכוי שקריאה שגויה אחת תכשיל את המשימה כולה.
AI, machine learning and LLM insights.
לולאת ה-ReAct מאפשרת למודל לבחור את הצעדים שלו, אך כל מחזור של מחשבה-פעולה-תצפית מוסיף עלות הסקה נוספת, מה שמגדיל את השיהוי ומעלה את הסיכוי שקריאה שגויה אחת תכשיל את המשימה כולה.
ה-Swarm מחליף מפקח יחיד ברשת שטוחה של סוכנים שווים המבקרים זה את זה ללא הרף, אך כל חילופי מידע מפעילים קריאה נפרדת למודל, מה שמנפח את עלויות המחשוב ואת השיהוי באופן דרמטי.
The agent rewrites queries, decides if a search is needed, breaks complex questions into sub-steps and self-corrects when evidence is weak, delivering citations for every claim and slashing token usage dramatically.
The opt-in service gives a small group of U.S. creators a dashboard that lists any video or image the AI flags as matching their verified facial features, after they clear a high-friction selfie-and-ID check via Jumio.
באמצעות החלפת חיפושי URL ב-hashes מבוססי תוכן, ה-COS API מאפשר לכל אתר לבקש קובץ שערך ה-SHA-256 שלו כבר ידוע לו, ובכך מאפשר לדפדפנים להגיש את אותו מודל AI בנפח 33 GB למקורות שונים מבלי להוריד אותו מחדש.
The revamped workflow scrapes queries with >500 impressions, <50 clicks, <1% CTR and position >20, then uses Llama 3 for headlines and Claude 3.5 Sonnet for full drafts, keeping each article under five cents while delivering a 15% lift in organic impressions.
מערכת ההערכה שלך תשקר לך לפני שהמודל שלך יעשה זאת. לוח התוצאות של ההערכה שלך הראה ששני המנועים נכשלו. Llama3.2 נכשל ב-5 מתוך 6 מקרים. Anthropic Sonnet נכשל ב-6 מתוך 6 מקרים...
הארכיטקטורה החדשה מעבדת שרשראות שלמות של אלף חומצות אמינו במעבר יחיד, תוך שימור אינטראקציות ארוכות טווח, מה שהופך את האימון על מאגרי רצפים עצומים למהיר וזול יותר.
The new x402 protocol embeds a payment token in each HTTP request, letting AI agents automatically deduct fees for data, compute or API calls. AWS’s Bedrock integration, called AgentCore Payments, adds built-in wallets and spend caps for seamless machine commerce.
WebMCP מחליף את ה-DOM-scraping השברירי במניפסט של קריאות כלים מוגדרות טיפוסית שרכיבי React רושמים בזמן ה-mount, ובכך מספק אינטראקציות מיידיות שאינן תלויות בפריסה (layout) עבור סוכני AI בכרום 149 ומעלה.
Backed by Sequoia, Felicis, Optum Ventures and Y Combinator, Bunkerhill plans pilot rollouts of its Carebricks agents—software that can pull EHR data, order labs and update care plans—aiming to relieve clinicians of routine coordination tasks.
מסגרת העבודה SEED מוסיפה שלב של עיבוד לאחר הפרק, שבו הסוכן קורא את התמליל של עצמו, כותב שיעורים בשפה טבעית ומזקק אותם לעדכוני מדיניות — ובכך הופך אותות תגמול דלילים לאות אימון עשיר וניתן לשימוש חוזר.
MCP version 2, live July 28 2026, removes the initialize handshake, the Mcp-Session-Id header and three legacy subsystems, making the protocol fully stateless. Teams can now run on autoscaled or serverless pods without session-affinity constraints.
ה-Inferentia2 INF2.24xlarge שחזר בדיוק את זרם הטוקנים של ייחוס ה-CPU, אך בשתי ההרצות הוזן פרומפט לא תקין שחסר בו תבנית צ'אט וסימני תור, מה ששלח את Gemma-4 ללולאה אינסופית של שטויות. החומרה רק שיקפה באג בקוד הייחוס.
By adding a PostgreSQL logging layer that captured model name, token count, task type, and per-call cost, the researcher identified that raw SEC search results were inflating prompts. Summarizing those results and routing simple checks to a cheaper model drove the 31% savings and nudged accuracy up
The loan uses SambaNova’s SN50 inference chips as collateral, a first-of-its-kind asset class that promises up to 16x faster token processing, lower power draw and simpler cooling than traditional GPU farms.
טורבאלדס אמר לרשימת התפוצה שכל מי שמתנגד לכלי בינה מלאכותית כמו ה-Sashiko reviewer החדש יכול פשוט לפצל את הליבה (fork), ובכך מדגיש את עמדתו לפיה יתרון טכני גובר על התנגדות אידיאולוגית.
The repo ships the Rust source for the CLI, terminal UI, and runtime, letting you examine how context is assembled, tools are orchestrated, and edits are managed – and even replace the underlying model without breaking the workflow.
Moonshot AI’s Kimi K3, a 2.8-trillion-parameter open-weight model, outscored GPT-5.6 on the real-world AA-Briefcase benchmark and is now available via API, with weights slated for public release on July 27.
The latest versions add task orchestration, approval workflows, testing rigs and UI-aware browsing, letting an AI pick tickets, run autonomous agents, get human sign-off and ship visual previews—all from a single console.
The service, available only to US-based developers, bills $1.25 for every million input tokens and $4.25 for every million output tokens, with a $20 credit for new accounts and a waitlist for international users.
The Sepsis Model That Learned To Predict The Doctor In 2021, Michigan Medicine tested the Epic Sepsis Model. They looked at 38,455 hospital stays. Epic claimed its accuracy was hi…
PrismML’s Bonsai 27B shrinks from a 54.7 GB full-precision model to a 3.9 GB 1-bit file, fitting on a flagship iPhone while retaining 89.5% of its quality. The claim relies on Apple’s MLX stack and delivers about 11 tokens per second, but memory, heat and battery impacts remain unmeasured.
מודל ה-GPT-Red הפנימי של OpenAI הפחית את הכשלים בסדרת GPT-5.6 Sol פי שש, אך נייר המחקר מציע רק תיאוריה ולא כלי עזר להורדה, מה שמותיר לצוותים קטנים לבנות בדיקות דטרמיניסטיות משלהם.