রিয়েল-টাইম কোলাবরেশন দেখতে খুব সহজ মনে হয়, যতক্ষণ না আপনি এর পেছনের জটিলতাগুলো দেখতে পাচ্ছেন। একজন টাইপ করছেন। অন্যজন তিন প্যারাগ্রাফ উপরে একটি লাইন মুছে ফেললেন। তৃতীয় একজন Stack Overflow থেকে একটি স্নিপেট পেস্ট করলেন। কোনোভাবে ডকুমেন্টটি একটি একক, সুসংগত অবস্থায় পৌঁছে যায়। WebSockets বা distributed state সম্পর্কে পূর্ব অভিজ্ঞতা না থাকা সত্ত্বেও একদম শুরু থেকে এই সাবলীলতা তৈরি করাটা অনেকটা অবিবেচকের মতো মনে হতে পারে। তবে এটিই আসলে শেখার সঠিক উপায় বলে মনে হয়।

এই প্রজেক্টটি একদম শূন্য থেকে শুরু। কোনো ধার করা বয়লারপ্লেট (boilerplate) নেই। কোনো পালিশ করা ইউটিউব ওয়াকথ্রু (walkthrough) নেই যেখানে কঠিন অংশগুলো ৩০ সেকেন্ডের একটি মন্টেজে বাদ দিয়ে দেওয়া হয়। লক্ষ্য হলো একটি কোলাবোরেটিভ কোড এডিটর তৈরি করা যেখানে একাধিক ব্যবহারকারী একই সাথে একই ফাইল এডিট করতে পারবেন এবং একে অপরের পরিবর্তন—এবং একে অপরের কার্সার—তাতে সাথে সাথে দেখতে পাবেন। সেখানে পৌঁছানোর জন্য ট্রান্সপোর্ট লেয়ার, কনসিস্টেন্সি মডেল এবং ডকুমেন্ট নষ্ট না করে কনকারেন্ট এডিটগুলো মার্জ করার কঠিন সমস্যাটি সমাধান করতে হবে।

"রিয়েল-টাইম" আসলে বলতে কী বোঝায়

বেশিরভাগ ওয়েব অ্যাপ্লিকেশন রিকোয়েস্ট-রেসপন্স সাইকেলে অভ্যস্ত। আপনি একটি ফর্ম জমা দেন, সার্ভার তা সেভ করে, আপনি পেজটি রিফ্রেশ করেন। রিয়েল-টাইম কোলাবরেশন এই নিয়মটি সম্পূর্ণ ভেঙে দেয়। প্রতিটি কি-স্ট্রোক (keystroke) হলো একটি ইভেন্ট যা প্রতিটি সংযুক্ত ক্লায়েন্টে মিলি-সেকেন্ডের মধ্যে ছড়িয়ে পড়তে হবে এবং এমন একটি ক্রমে পৌঁছাতে হবে যা তথ্যের অর্থ বজায় রাখে।

WebSockets এখানে ট্রান্সপোর্টের জন্য সবচেয়ে উপযুক্ত পছন্দ কারণ এটি ক্লায়েন্ট এবং সার্ভারের মধ্যে একটি স্থায়ী, ফুল-ডুপ্লেক্স কানেকশন বজায় রাখে। HTTP পোলিংয়ের মতো নয়, যা প্রতি কয়েক সেকেন্ড অন্তর "নতুন কিছু আছে কি?" জিজ্ঞাসা করে ব্যান্ডউইথ নষ্ট করে, একটি WebSocket খোলা থাকে। যখন ব্যবহারকারী A একটি সেমিকোলন টাইপ করেন, সেই ক্যারেক্টারটি একটি মেসেজ হিসেবে সকেটের মাধ্যমে একটি কেন্দ্রীয় সার্ভারে যায় এবং তারপর ব্যবহারকারী B এবং C-এর কাছে পৌঁছে যায়। এই অংশটি তুলনামূলকভাবে সহজ।

কঠিন অংশটি হলো যখন B এবং C ঠিক একই মুহূর্তে টাইপ করেন তখন কী ঘটে। যদি উভয় পরিবর্তন প্রায় একই সাথে সার্ভারে পৌঁছায়, তবে কোনটি কার্যকর হবে? আপনি যদি কেবল আগমনের ক্রম অনুযায়ী মেসেজগুলো ব্রডকাস্ট করেন, তবে ক্যারেক্টার হারিয়ে যাওয়া বা টেক্সট এলোমেলো হয়ে যাওয়ার ঝুঁকি থাকে। সাধারণ "লাস্ট-রাইট-উইন্স" (last-write-wins) স্ট্র্যাটেজিগুলো ব্যর্থ হয় কারণ সেগুলো ব্যবহারকারীর উদ্দেশ্যকে (intent) উপেক্ষা করে। আমি যদি লাইনের শুরুতে "hello" টাইপ করি আর আপনি যদি লাইনের শুরুতে "world" টাইপ করেন, তবে ফলাফল এমন হওয়া উচিত নয় যেখানে আমাদের মধ্যে একজনের লেখা মুছে যায়। এটি ডিটারমিনিস্টিকভাবে "helloworld" বা "worldhello" হওয়া উচিত। এটি অর্জন করার জন্য এমন একটি সিনক্রোনাইজেশন স্ট্র্যাটেজি প্রয়োজন যা ডকুমেন্টের গঠন বুঝতে পারে।

কেন একদম শূন্য থেকে শুরু করা গুরুত্বপূর্ণ

এমন চমৎকার সব ফ্রেমওয়ার্ক রয়েছে যা এই জটিলতাগুলো লুকিয়ে রাখে। Yjs, Automerge এবং Socket.IO এই কষ্টগুলো এড়িয়ে একটি কার্যকরী প্রোটোটাইপ এক বিকেলেই তৈরি করে দিতে পারে। কিন্তু এর নিচের মৌলিক বিষয়গুলো (primitives) না বুঝে এগুলো ব্যবহার করা অনেকটা ইনস্ট্রুমেন্ট বা যন্ত্রপাতির কাজ না জেনে অটোপাইলটে বিমান চালানোর মতো। যখন টার্বুলেন্স বা অস্থিরতা দেখা দেয়—এবং ডিস্ট্রিবিউটেড সিস্টেমে এটি সবসময়ই ঘটে—তখন আপনার জানতে হবে সমস্যাটি কি আপনার নেটওয়ার্ক লেয়ারে, কনফ্লিক্ট রেজোলিউশনে নাকি আপনার ডেটা মডেলে।

এখানে প্রতিশ্রুতি হলো লাইব্রেরির ওপর নির্ভর করার আগে ধারণাগুলো শেখা। এর মানে হলো ম্যানুয়ালি চিন্তা করা যে নিচের ঘটনাগুলো ঘটলে কী হবে:

  • একজন ক্লায়েন্ট টাইপ করার মাঝপথে ডিসকানেক্ট হয়ে গেলেন এবং দশ সেকেন্ড পরে আবার রিকানেক্ট হলেন
  • দুইজন ব্যবহারকারী একই সাথে একই কার্সার পজিশনে টেক্সট ইনসার্ট করছেন
  • একজন ব্যবহারকারী এমন একটি ব্লক মুছে ফেলছেন যা অন্য একজন ব্যবহারকারী সক্রিয়ভাবে এডিট করছেন
  • সার্ভার ক্র্যাশ করল এবং একটি নতুন নোডকে একদম শুরু থেকে ডকুমেন্টের স্টেট পুনর্গঠন করতে হচ্ছে

এই সমস্যাগুলোর সমাধানের জন্য Operational Transformation (OT) এবং Conflict-free Replicated Data Types (CRDTs) হলো দুটি প্রধান পদ্ধতি। Google Docs তার প্রাথমিক আর্কিটেকচার OT-এর ওপর ভিত্তি করে তৈরি করেছিল, যেখানে অপারেশনগুলো প্রয়োগ করার আগে একটি কেন্দ্রীয় সার্ভারের প্রয়োজন হয় যাতে তারা একে অপরের বিপরীতে অপারেশনগুলো পরিবর্তন করতে পারে। বিপরীতে, CRDTs এমনভাবে ডিজাইন করা হয়েছে যাতে কোনো সমন্বয় ছাড়াই লোকালি কনকারেন্ট আপডেটগুলো মার্জ করা যায়, যা তাদের পিয়ার-টু-পিয়ার বা এজ-ভিত্তিক সেটআপের জন্য আকর্ষণীয় করে তোলে। এগুলোর মধ্যে একটি বেছে নেওয়া—বা হাইব্রিড পদ্ধতি ব্যবহার করা—এর জন্য মেমরি ব্যবহার, কনভারজেন্স গ্যারান্টি এবং ইমপ্লিমেন্টেশন জটিলতার মধ্যে ভারসাম্য বোঝা প্রয়োজন। শুধু এই বিষয়ে পড়া যথেষ্ট নয়; পরিকল্পনা হলো উভয় পদ্ধতির সাধারণ (naive) এবং উন্নত (refined) সংস্করণ ইমপ্লিমেন্ট করা যাতে দেখা যায় কোথায় সেগুলো ব্যর্থ হয়।

পুনর্নির্মাণ, ভুল এবং অচল পথসমূহ

Expectations are calibrated honestly. There will be stretches where nothing works. A first attempt might use simple JSON patches to represent text changes, only to discover that JSON has no concept of “index 5 in a paragraph,” so two concurrent insertions at the same index overwrite each other instead of merging. A second attempt might build a custom linear history log, only to realize that replaying that log is Big O nightmare when the document grows. A third attempt might get WebSockets working locally, then fall apart over a real network where packet loss and variable latency rewrite the rules.

That friction is the point. Copying a working repository would skip the investigation of why the queue flushes in that particular order, or why the server maintains a version vector. Rebuilding the same component three times is slow, but it forces an understanding of the boundary between what the framework does and what your own logic must handle.

The documentation of this process will not be a highlight reel. It will include the wrong turns. For example, building presence awareness—knowing who is online and where their cursor is—seems like a cosmetic feature until you realize it depends on the same consistency model as the text itself. If user A sees user B’s cursor at column 10, then user B inserts four characters, where does that cursor move? Without a shared understanding of document topology, presence data drifts from reality. Solving that requires coupling the cursor position to the underlying data structure’s identity, not just its numerical index. These are the kinds of details that tutorials skim over because they are tedious, not because they are unimportant.

What Comes Next

The immediate roadmap is sparse by design. The first milestones will be:

  • A raw WebSocket server that echoes character events, to feel the latency and connection lifecycle firsthand
  • A simple string buffer on the client to understand why naive insertion ordering fails under concurrency
  • A from-scratch CRDT for ordered sequences, however inefficient, to see the commutative property in action
  • Gradual integration with an actual code editor surface, likely something like CodeMirror or Monaco, to grapple with the mismatch between the editor’s imperative API and the functional nature of operational history

Each step will come with a written rationale. Why this approach and not that one? What assumptions were disproven? What abstraction leaked?

A Real Takeaway

Starting a project like this without experience in WebSockets or CRDTs is intimidating, but expertise is often just repeated confusion with better labels. The objective is not a fast finish. It is a system whose behavior is predictable because every layer was built with intent rather than imported with hope.

If you have built collaborative software before—whether a text editor, a design tool, or a game state sync engine—share the failure modes that caught you off guard. If you are learning these systems too, follow along. The code will arrive slowly, and it will be rewritten often. Day 0 begins now.