درک ویدئو همزمان دو پرسش دشوار را مطرح می‌کند: چه چیزی در کادر است و به کجا می‌رود؟ سیستم‌های بینایی ماشین به‌طور سنتی به این پرسش‌ها یکی پس از دیگری پاسخ می‌دهند. ابتدا آن‌ها قطعه‌بندی (segmentation) می‌کنند و اشیاء را از میان پیکسل‌ها جدا می‌سازند. سپس آن‌ها را ردیابی (tracking) می‌کنند و آن اشیاء را در طول زمان به هم متصل می‌نمایند. این تفکیک باعث ایجاد اصطکاک می‌شود. خطاهای مرحله اول به مرحله دوم سرایت می‌کنند. مرزها جابه‌جا می‌شوند. هویت‌ها تغییر می‌کنند. وقتی افراد یا وسایل نقلیه روی هم قرار می‌گیرند، کل خط لوله (pipeline) متوقف می‌شود.

کارهای اخیر روی فرمول‌بندی‌های multi-cut مسیر متفاوتی را ارائه می‌دهند. این روش به‌جای زنجیره‌ای کردن دو مدل، قطعه‌بندی و ردیابی را به‌عنوان یک مسئله بهینه‌سازی واحد در نظر می‌گیرد. نتیجه آن مرزهای دقیق‌تر، تغییر هویت‌های کمتر و خط لوله‌ای است که در صحنه‌های شلوغ، یکپارچه باقی می‌ماند.

مشکل خط لوله‌ها (Pipelines)

اکثر سیستم‌های عملیاتی ابتدا تشخیص (detection) یا قطعه‌بندی را اجرا می‌کنند و سپس خروجی را به یک ماژول ردیابی تحویل می‌دهند. همین مرحله‌ی تحویل است که باعث بروز خطا می‌شود. یک مدل قطعه‌بندی که روی تصاویر ثابت آموزش دیده است، ممکن است در فریم دهم ماسکی تولید کند که کمی بزرگ باشد و در فریم یازدهم ماسکی که کمی کوچک باشد. یک ردیاب که فقط به مرکز باکس‌ها یا فاصله‌های embedding نگاه می‌کند، باید تصمیم بگیرد که آیا آن دو ماسک متعلق به یک شیء هستند یا خیر. این ردیاب راهی ندارد که بازخورد دهد و به مدل قطعه‌بندی بگوید که مرز اشتباه به نظر می‌رسد. این دو سیستم با هم ارتباط برقرار نمی‌کنند.

این جداسازی زمانی که اشیاء با هم تعامل دارند، هزینه‌بر می‌شود. یک تقاطع خیابانی را تصور کنید که در آن عابران پیاده از میان یکدیگر عبور می‌کنند، یا یک بازوی رباتیک که برای برداشتن یک قطعه خاص، از میان پشته‌ای از جعبه‌ها عبور می‌کند. در این لحظات، ردیاب‌های سنتی برای حفظ هویت، به مدل‌های حرکتی یا ویژگی‌های ظاهری تکیه می‌کنند. وقتی انسداد (occlusion) بیش از چند فریم طول می‌کشد، این نشانه‌ها از کار می‌افتند. ردیاب یا یک هویت جدید برای همان شیء می‌سازد یا هویت را به شیء دیگری نسبت می‌دهد. در همین حال، مدل قطعه‌بندی به تولید ماسک‌ها ادامه می‌دهد، در حالی که کاملاً بی‌خبر است که برچسب‌ها دچار انحراف شده‌اند. پاکسازی این انحراف مستلزم پس‌پردازش سنگین یا برچسب‌گذاری دستی است که هدف خودکارسازی را از بین می‌برد.

مدل‌های سنتی اغلب دقیقاً زمانی که اشیاء روی هم قرار می‌گیرند، ردیابی را از دست می‌دهند. این خطاها تصادفی نیستند؛ بلکه ساختاری هستند. خط لوله‌ای که با زمان به‌عنوان یک موضوع ثانویه برخورد می‌کند، ناگزیر با تداوم (continuity) دچار مشکل خواهد شد.

آنچه Multi-Cut ارائه می‌دهد

یک فرمول‌بندی multi-cut دیوار میان قطعه‌بندی و ردیابی را از بین می‌برد. این روش به‌جای تولید توالی‌ای از ماسک‌های جدا از هم و سپس دوختن آن‌ها به یکدیگر، گرافی می‌سازد که هم فضا و هم زمان را در بر می‌گیرد. هر ناحیه بالقوه از یک شیء در هر فریم به یک گره (node) تبدیل می‌شود. یال‌ها (edges) نواحی را که ممکن است متعلق به یک موجودیت واحد باشند، هم در یک فریم و هم در فریم‌های متوالی، به هم متصل می‌کنند. سپس الگوریتم، روش بهینه برای بریدن آن یال‌ها را می‌یابد، به‌گونه‌ای که مولفه‌های متصلِ باقی‌مانده، اشیاء منسجمی را تشکیل دهند که به‌طور طبیعی در ویدئو حرکت

Crowded scenes. Surveillance and crowd-counting applications lose accuracy when people bunch together. Separate trackers often merge identities or create phantom objects in the gaps between bodies. By linking visual evidence directly into the tracking logic, the multi-cut approach maintains object identity better in crowded scenes. Individuals stay distinct even when their bounding boxes almost completely overlap.

Boundary integrity. In robotics, a pick-and-place system needs exact object contours to plan grasps. A segmentation model that ignores temporal context might include part of the background or clip off a corner of the item. When segmentation and tracking share a single objective, the model learns that boundaries should be temporally stable. The resulting masks snap more cleanly to real edges, which means fewer failed grasps and less need for conservative safety margins.

Building Better Pipelines

For engineers working in video analysis or robotics, this shift has concrete implications for how you architect your system.

Stop thinking of detection, segmentation, and tracking as three separate microservices that pass tensors down a conveyor belt. Look instead for frameworks that expose a joint objective. If you are building a custom pipeline, consider whether your graph structure can encode both spatial affinity and temporal continuity in the same loss. Even if you do not implement the full multi-cut solver yourself, the principle still applies. Add constraints that tie appearance over time back into the quality of the per-frame mask.

Test aggressively on occlusion. A demo video with isolated objects moving in simple patterns will not reveal the weaknesses of a pipelined approach. Collect sequences where targets overlap, where lighting changes mid-occlusion, and where objects re-enter from off-screen. These are the moments that separate a glued-together pipeline from a truly unified one.

Prepare your annotation strategy carefully. Joint models often need different ground truth structures than pure segmentation or pure tracking datasets. You may need to label instance identities consistently across frames, not just pixel classes per image. Investing in that labeling upfront pays off later in reduced manual correction during deployment.

The Real Takeaway

Treating segmentation and tracking as independent chores made sense when compute and model capacity were scarce. It no longer does. A multi-cut formulation shows that the two tasks are really one: finding coherent objects that persist through time. By solving them together, you get masks that respect motion and tracks that respect shape. The result is a video understanding pipeline that reduces errors between frames, creates cleaner boundaries around moving objects, and stays accurate through the messy reality of occlusions and crowds.