Every few months the open-source community mints another AI framework. Most of them wrap Python bindings around heavy C++ kernels, or they stack abstraction layers so high that the runtime alone weighs more than the models they serve. CatAI moves in the opposite direction. It is a native AI engine written entirely in C++, built from the tensor math upward. The point is not to create yet another friendly skin over PyTorch. The point is to own every byte of memory and every cycle of compute, starting at the hardware boundary.
Why Another Engine?
If you have shipped anything to production, you already know the pain. Pull a standard deep-learning stack into a container and watch the image bloat to multiple gigabytes. Dependencies fight each other. The Python interpreter adds latency. The dispatcher that routes ops to CUDA or CPU introduces subtle overhead that becomes impossible to profile once it disappears into a dozen nested frameworks. For edge devices, embedded robotics, or latency-sensitive backends, that tax is real. A pure C++ engine eliminates the middleman. It talks to the operating system and the silicon directly, with no garbage collection, no global interpreter lock, and no serialization dance between languages.
CatAI treats this as a feature, not a compromise. The project is being written from scratch in C++ because the author wants to decide exactly how tensors live in RAM, how they move through cache hierarchies, and how kernels are scheduled across threads. That is not masochism. It is the only way to guarantee that behavior is predictable when you are squeezing performance out of limited hardware.
What “From Scratch” Actually Means
In most modern frameworks, tensor math is handled by opaque calls into vendor libraries like cuDNN, oneMKL, or MPS. That is perfectly sensible for shipping fast, but it hides the mechanics of the operation. CatAI is writing its own core tensor math and memory layouts. That means designing the fundamental data structures that hold multi-dimensional arrays, choosing how strides and offsets are calculated, and deciding whether to store data in row-major, column-major, or custom tiled formats depending on the access pattern.
This is deep systems work. When you write a matrix-multiply kernel by hand, you stop thinking in terms of torch.matmul and start thinking about L1 cache lines, register pressure, and loop tiling. You decide whether to block for 32x32 tiles or 64x64 based on the SIMD width of the target CPU. You align allocations to 64-byte boundaries so AVX-512 loads do not cross cache lines. You question whether std::vector is the right container for tensor storage, or whether a custom arena allocator gives you better locality and zero fragmentation across an entire inference graph.
Memory layout is equally critical. A naive naïve n-dimensional array can kill performance if the channels-last image data is accessed in a channels-first pattern. In CatAI, these layouts are first-class citizens, not afterthoughts handled by a graph optimizer running at export time.
The Optimization Mindset
Bare-metal optimization sounds like a buzzword until you start counting nanoseconds. It means fusing operations so intermediate results never leave the CPU registers or L1 cache. It means implementing a layer-norm followed by a GELU as a single kernel, saving an entire round-trip to DRAM. It means writing your own thread pool instead of leaning on OpenMP defaults, because you know your workload is bursty and you do not want the runtime spawning and joining threads every forward pass.
It also means understanding when not to write assembly. Sometimes the compiler vectorizes a loop better than hand-written intrinsics. The discipline ismeasurement: profile, hypothesize, change one variable, and profile again. This engine is being built by people who enjoy that grind. If you have ever spent an afternoon rewriting a convolution loop to shave two milliseconds off a batch, you already understand the culture.
Who We Need
This is not a one-person show. Building a backend from zero requires distinct skills that rarely overlap in a single brain. If you are reading this and considering whether to jump in, here is where you might fit:
C++ geliştiricileri; modern standartları bilen ancak şablonların (templates) ne zaman derleme şişkinliğine (compilation bloat) yol açacağını da anlayan kişiler. Gerektiğinde ham işaretçiler (raw pointers), uygun olduğunda ise akıllı işaretçiler (smart pointers) konusunda rahat olmalı ve sözdizimi kolaylığından (syntax sugar) en az ikili dosya boyutu (binary size) kadar önem vermelisiniz.
Matematik uzmanları; standart olmayan aktivasyonlar için geri geçiş (backward-pass) gradyanlarını türetebilen, karma hassasiyetli (mixed-precision) eğitimde sayısal kararlılık üzerine akıl yürütebilen ve algoritmaları kod haline gelmeden önce optimize edebilen kişiler. Eğer bir log-sum-exp hilesinin neden önemli olduğunu açıklayabiliyorsanız, doğru zihinsel alandasınız demektir.
Düşük seviyeli bellek uzmanları; tahsis ediciler (allocators), sayfa hataları (page faults) ve NUMA topolojisi üzerine düşünen kişiler. Motorun; grafik yürütme için bellek havuzlarına (memory pools), çekirdekler (kernels) için geçici tamponlara (scratch buffers) ve sızıntı veya parçalanma (fragmentation) olmadan eğitim adımları boyunca tensör depolamasını yeniden kullanma stratejilerine ihtiyacı var.
Sistem mühendisleri; yanlış yerleştirilmiş bir sistem çağrısının (syscall) tüm bir eğitim döngüsünü nasıl durdurabileceğini anlayan kişiler. Zamanlama (scheduling), G/Ç (I/O) ve senkronizasyon ilkeleri, matematiği bir arada tutan yapıştırıcıdır.
Bu dört alanın hepsinde dünya çapında bir uzman olmanıza gerek yok. Çoğu katkıda bulunan, tek bir çekirdek (kernel) veya tek bir tahsis edici (allocator) sahiplenerek başlayacak ve mimari sağlamlaştıkça geri kalanını öğrenecektir.
Mimari ve Özel Matematik
Arka uç mantığı iş birliği içinde inşa ediliyor ve bu, mimari tartışmalarıyla başlıyor. Motor, tüm modelin çalışma zamanından önce tanımlandığı ve optimize edildiği statik bir hesaplama grafiği mi kullanacak? Yoksa otomatik türev için bir bant (tape) ile anlık yürütmeyi (eager execution) mi destekleyecek? Otomatik türev (autodiff) nasıl temsil edilecek—operatör aşırı yüklemesi (operator overloading), kaynak dönüşümü (source transformation) veya bir grafik IR mi? Bu kararlar geri kalan her şeyi şekillendiriyor.
Özel sinir ağı matematiği, standart katmanları yeniden uygulamaktan daha fazlasını ifade eder. Bu, yenilerini icat etme özgürlüğü demektir. Standart olmayan seyrek bir çekirdeğe (sparse kernel) sahip bir konvolüsyon varyantı veya literatürde adı olmayan bir aktivasyon fonksiyonu istiyorsanız, C++ ileri (forward) ve geri (backward) geçişlerini yazar ve bunları doğrudan motora bağlarsınız. Savaşmanız gereken bir Python API'si veya gerek duyulan bir monkey-patching yoktur. Matematik koddur ve kod arayüzdür.
Nasıl Katılabilirsiniz
Eğer bu size hitap ediyorsa, projenin tam dökümü ve mevcut yol haritası yazarın Dev.to gönderisinde ayrıntılı olarak belgelenmiştir. Detayları okuyabilir, şimdiye kadar nelerin inşa edildiğini görebilir ve tam olarak nerede yardıma ihtiyaç duyulduğunu anlayabilirsiniz.
Proje detayları: https://dev.to/banana_cool/building-a-native-c-ai-engine-catai-from-scratch-looking-for-collaborators-l8m
Hemen bir pull request göndermeye taahhüt vermeden takılmak, soru sormak veya ilerlemeyi takip etmek isteyenler için bir Telegram grubu da bulunmaktadır.
Topluluk: https://t.me/GyaanSetuAi
Asıl Çıkarım
Modern yapay zeka yığını bir kara kutu haline geldi. Çerçeveleri (frameworks) sihirli cihazlar gibi ele alıyoruz: veri giriyor, model çıkıyor ve biz bu kapalılığın dağıtım (deployment) aşamasında bizi vurmamasını umuyoruz. CatAI bu konforu reddeder. Bu şekilde inşa etmek daha yavaştır. Daha fazla kod yazacak, daha fazla segfault ayıklayacak ve üst düzey çerçevelerin sizden gizlediği varsayımları yeniden düşüneceksiniz. Ancak makinenin neden bu şekilde davrandığını da anlayacaksınız. Herkesin donanımı soyutlamak için yarıştığı bir endüstride, ters yöne gitmenin ve metale dokunmanın (touching the metal) gerçek bir değeri vardır. Bu anlayış, API çağıran biriyle sistem inşa eden birini birbirinden ayıran şeydir.
