ਸਥਾਨਕ (local) ਵੱਡੇ ਭਾਸ਼ਾ ਮਾਡਲਾਂ (large language models) ਨੂੰ ਚਲਾਉਣਾ ਅਕਸਰ ਤੁਹਾਨੂੰ ਹਾਰਡਵੇਅਰ ਦੀਆਂ ਸੀਮਾਵਾਂ ਵਿੱਚ ਫਸਾ ਦਿੰਦਾ ਹੈ। NVIDIA ਵਰਤੋਂਕਾਰ CUDA ecosystem ਦੇ ਅੰਦਰ ਰਹਿੰਦੇ ਹਨ। Apple ਡਿਵੈਲਪਰ Metal ਦੀ ਵਰਤੋਂ ਕਰਦੇ ਹਨ। ਬਾਕੀ ਸਭ ਉਮੀਦ ਕਰਦੇ ਹਨ ਕਿ ਉਨ੍ਹਾਂ ਦਾ GPU OpenCL ਨੂੰ ਸਮਝੇ ਜਾਂ ਸਿਰਫ਼ CPU 'ਤੇ ਨਿਰਭਰ ਰਹੇ। ਇਹ ਟੁਕੜੇਬੰਦੀ (fragmentation) ਡੈਸਕਟਾਪ AI ਐਪਲੀਕੇਸ਼ਨ ਨੂੰ ਲਾਂਚ ਕਰਨਾ ਉਨਾ ਹੀ ਮੁਸ਼ਕਲ ਬਣਾ ਦਿੰਦੀ ਹੈ ਜਿੰਨਾ ਕਿ ਹੋਣਾ ਚਾਹੀਦਾ ਹੈ। TensorSharp ਨੇ ਹੁਣ ਇੱਕ Vulkan backend ਜੋੜ ਕੇ ਇਸ ਸਮੱਸਿਆ ਨੂੰ ਹੱਲ ਕਰਨ ਦੀ ਕੋਸ਼ਿਸ਼ ਕੀਤੀ ਹੈ, ਜਿਸ ਨਾਲ ਇੰਜਣ ਨੂੰ ਵੱਖ-ਵੱਖ ਵੈਂਡਰਾਂ ਦੇ discrete GPUs 'ਤੇ ਚੱਲਣ ਦਾ ਇੱਕ ਭਰੋਸੇਯੋਗ ਰਸਤਾ ਮਿਲ ਗਿਆ ਹੈ।

Vulkan ਕਿਉਂ ਸਥਿਤੀ ਬਦਲ ਦਿੰਦਾ ਹੈ

Vulkan ਬਾਰੇ ਆਮ ਤੌਰ 'ਤੇ ਗੇਮਿੰਗ ਸਰਕਲਾਂ ਵਿੱਚ ਚਰਚਾ ਕੀਤੀ ਜਾਂਦੀ ਹੈ, ਪਰ ਇੱਕ low-overhead, cross-platform compute API ਵਜੋਂ, ਇਹ inference ਲਈ ਵੀ ਉਨਾ ਹੀ ਮਹੱਤਵਪੂਰਨ ਹੈ। ਇਹ ਉਸ ਹਾਰਡਵੇਅਰ ਤੱਕ ਪਹੁੰਚਦਾ ਹੈ ਜਿਸ ਨੂੰ CUDA ਨਜ਼ਰਅੰਦਾਜ਼ ਕਰ ਦਿੰਦਾ ਹੈ। ਜਿਵੇਂ ਕਿ Intel UHD ਅਤੇ Iris Xe integrated chips, ਪੁਰਾਣੇ discrete cards, ਅਤੇ ਬਿਨਾਂ NVIDIA ਸਟਿੱਕਰ ਵਾਲੇ ਬਜਟ Windows ਲੈਪਟਾਪ। ਇੱਕ local inference engine ਲਈ, ਇਹ ਪਹੁੰਚ ਇੱਕ ਵਿਹਾਰਕ ਸ਼ਕਤੀ ਹੈ। ਇੱਕ ਡਿਵੈਲਪਰ ਇੱਕ ਸਿੰਗਲ binary path ਲਾਂਚ ਕਰ ਸਕਦਾ ਹੈ ਜੋ CUDA-only ਹੱਲ ਦੇ ਮੁਕਾਬਲੇ ਕਿਤੇ ਜ਼ਿਆਦਾ ਮਸ਼ੀਨਾਂ 'ਤੇ ਚੱਲ ਸਕਦਾ ਹੈ।

TensorSharp ਦਾ Vulkan support GGML project ਰਾਹੀਂ ਸ਼ੁਰੂ ਹੋਇਆ। ਇਹ integration ਅੱਜ ਕੰਮ ਕਰ ਰਿਹਾ ਹੈ, ਹਾਲਾਂਕਿ ਲੇਖਕ ਬਾਅਦ ਵਿੱਚ ਇੱਕ native Vulkan backend ਬਣਾਉਣ ਦੀ ਯੋਜਨਾ ਬਣਾ ਰਿਹਾ ਹੈ। GGML ਨੂੰ ਇੱਕ ਪੁਲ (bridge) ਵਜੋਂ ਵਰਤਣਾ ਇੱਕ ਸਹੀ ਕਦਮ ਸੀ। ਇਹ ਆਰਕੀਟੈਕਚਰ ਦੀ ਪੁਸ਼ਟੀ ਕਰਦਾ ਹੈ ਅਤੇ ਤੁਰੰਤ ਟੈਸਟਰਾਂ ਦੇ ਹੱਥਾਂ ਵਿੱਚ ਹਾਰਡਵੇਅਰ ਪਹੁੰਚਾਉਂਦਾ ਹੈ। ਅਬਸਟਰੈਕਸ਼ਨ (abstraction) ਦੇ ਬੋਝ ਨੂੰ ਘਟਾਉਣ ਅਤੇ C#-centric ਇੰਜਣ ਨੂੰ command buffers ਅਤੇ memory barriers 'ਤੇ ਬਿਹਤਰ ਕੰਟਰੋਲ ਦੇਣ ਲਈ ਇੱਕ native backend ਲਿਆਂਦਾ ਜਾਵੇਗਾ।

ਹੁਣ ਤੱਕ ਟੈਸਟਿੰਗ ਸਤ੍ਹਾ ਕਿਹੋ ਜਿਹੀ ਹੈ

ਵੈਲੀਡੇਸ਼ਨ (Validation) ਵਿੱਚ ਪਹਿਲਾਂ ਹੀ ਦੋ ਬਹੁਤ ਵੱਖਰੇ Windows ਕੌਂਫਿਗਰੇਸ਼ਨ ਸ਼ਾਮਲ ਹਨ। ਡਿਵੈਲਪਰ ਨੇ NVIDIA GeForce RTX 3080 Laptop GPU ਅਤੇ ਸਾਧਾਰਨ Intel UHD Graphics 'ਤੇ ਟੈਸਟ ਕੀਤਾ। ਦੋਵੇਂ ਵਧੀਆ ਚੱਲੇ। ਇਹ ਰੇਂਜ ਨੋਟ ਕਰਨ ਯੋਗ ਹੈ। Inference ਦੀ ਦੁਨੀਆ ਵਿੱਚ, discrete high-wattage silicon ਅਤੇ ਬੇਸਿਕ integrated graphics ਬਹੁਤ ਘੱਟ ਹੀ ਇੰਨੀ ਆਸਾਨੀ ਨਾਲ ਇੱਕੋ ਟੈਸਟ ਸਤ੍ਹਾ 'ਤੇ ਚੱਲਦੇ ਹਨ। ਜੇਕਰ ਤੁਸੀਂ ਬਿਨਾਂ ਕਿਸੇ dedicated GPU ਵਾਲਾ ਹਲਕਾ ਲੈਪਟਾਪ ਚਲਾ ਰਹੇ ਹੋ, ਤਾਂ TensorSharp ਹੁਣ ਇੱਕ ਅਸਲੀ acceleration path ਪ੍ਰਦਾਨ ਕਰਦਾ ਹੈ ਜੋ NVIDIA drivers 'ਤੇ ਨਿਰਭਰ ਨਹੀਂ ਕਰਦਾ।

ਇਸ ਮੈਟ੍ਰਿਕਸ ਵਿੱਚ ਕਮੀ AMD ਦੀ ਹੈ। ਅਜੇ ਤੱਕ ਕਿਸੇ ਵੀ Radeon hardware ਦਾ ਟੈਸਟ ਨਹੀਂ ਕੀਤਾ ਗਿਆ ਹੈ। ਜੇਕਰ ਤੁਹਾਡੇ ਕੋਲ AMD GPU ਹੈ, ਤਾਂ ਇਸ ਪ੍ਰੋਜੈਕਟ ਨੂੰ ਤੁਹਾਡੇ ਫੀਡਬੈਕ ਦੀ ਲੋੜ ਹੈ। RX 6000 ਜਾਂ 7000 ਸੀਰੀਜ਼ ਦੇ ਕਾਰਡਾਂ 'ਤੇ ਕਮਿਊਨਿਟੀ ਵੈਲੀਡੇਸ਼ਨ ਹੀ ਇੱਕ ਪ੍ਰਯੋਗਿਕ backend ਨੂੰ production-grade ਵਿਕਲਪ ਵਿੱਚ ਬਦਲਦੀ ਹੈ। ਜੇਕਰ ਇਹ ਖਰਾਬ ਹੁੰਦਾ ਹੈ ਤਾਂ 'issue' ਦਰਜ ਕਰੋ। ਜੇਕਰ ਇਹ ਵਧੀਆ ਕੰਮ ਕਰਦਾ ਹੈ ਤਾਂ ਵੀ ਦਰਜ ਕਰੋ। ਦੋਵੇਂ ਹੀ ਨਤੀਜੇ ਪ੍ਰੋਜੈਕਟ ਨੂੰ ਅੱਗੇ ਵਧਾਉਂਦੇ ਹਨ।

TensorSharp ਕੋਈ Wrapper ਨਹੀਂ ਹੈ

ਇਸ ਨੁਕਤੇ 'ਤੇ ਜ਼ੋਰ ਦੇਣ ਦੀ ਲੋੜ ਹੈ। TensorSharp, llama.cpp ਦੇ ਆਲੇ-ਦੁਆਲੇ ਕੋਈ C# binding ਨਹੀਂ ਹੈ। ਡਿਵੈਲਪਰ ਨੇ ਪੂਰਾ ਇੰਜਣ ਜ਼ਮੀਨ ਤੋਂ (from the ground up) ਬਣਾਇਆ ਹੈ। CPU backend ਸ਼ੁੱਧ C# ਹੈ। ਜਦੋਂ ਤੁਸੀਂ GPU ਤੋਂ ਬਿਨਾਂ inference ਚਲਾਉਂਦੇ ਹੋ, ਤਾਂ ਤੁਸੀਂ ਕਿਸੇ foreign function interface ਰਾਹੀਂ C++ binary ਵਿੱਚ ਜਾਣ ਦੀ ਬਜਾਏ managed code ਨੂੰ ਚਲਾ ਰਹੇ ਹੁੰਦੇ ਹੋ। ਇਹ ਪ੍ਰੋਜੈਕਟ CUDA, Apple ਦੇ MLX, ਅਤੇ GGML ਲਈ ਸਮਰਪਿਤ backends ਵੀ ਰੱਖਦਾ ਹੈ। ਉਸ ਆਰਕੀਟੈਕਚਰਲ ਸੁਤੰਤਰਤਾ ਦੇ ਬਾਵਜੂਦ, ਇਸਦੀ ਪ੍ਰਦਰਸ਼ਨ (performance) llama.cpp ਦੇ ਬਰਾਬਰ ਹੈ, ਜੋ ਕਿ ਉਹ ਰੈਫਰੈਂਸ ਪੁਆਇੰਟ ਹੈ ਜਿਸਦਾ ਮੋਸਟ local inference ਪ੍ਰੋਜੈਕਟ ਪਿੱਛਾ ਕਰਦੇ ਹਨ। ਇਹ ਬਰਾਬਰੀ ਬਹੁਤ ਮੁਸ਼ਕਲ ਨਾਲ ਪ੍ਰਾਪਤ ਕੀਤੀ ਗਈ ਹੈ। ਇਸਦਾ ਮਤਲਬ ਹੈ ਕਿ memory layout, kernel dispatch, ਅਤੇ tensor ops ਸਾਰੇ ਅਸਲ ਲੋਡ ਦੇ ਹੇਠਾਂ ਸਹੀ ਤਰ੍ਹਾਂ ਕੰਮ ਕਰਦੇ ਹਨ।

ਮਾਡਲ ਸਪੋਰਟ ਵਿੱਚ Gemma4, DiffusionGemma, ਅਤੇ Qwen3.6 ਸ਼ਾਮਲ ਹਨ। Runtime ਮਲਟੀਮੋਡਲ (multimodal) ਕੰਮ ਵੀ ਸੰਭਾਲਦਾ ਹੈ। Vision, audio, ਅਤੇ reasoning pipelines ਇੱਕੋ ਇੰਜਣ ਰਾਹੀਂ ਚੱਲਦੇ ਹਨ। ਜੇਕਰ ਤੁਸੀਂ ਇੱਕ ਅਜਿਹੇ ਡੈਸਕਟਾਪ ਸਹਾਇਕ (assistant) ਦਾ ਪ੍ਰੋਟੋਟਾਈਪ ਬਣਾ ਰਹੇ ਹੋ ਜੋ ਸਕ੍ਰੀਨਸ਼ੌਟ ਪੜ੍ਹਦਾ ਹੈ ਅਤੇ ਆਵਾਜ਼ ਦੇ ਨਿਰਦੇਸ਼ਾਂ ਨੂੰ ਸਵੀਕਾਰ ਕਰਦਾ ਹੈ, ਤਾਂ ਤੁਹਾਨੂੰ ਤਿੰਨ ਵੱਖ-ਵੱਖ runtimes ਨੂੰ ਜੋੜਨ ਅਤੇ ਇਹ ਪ੍ਰਾਰਥਨਾ ਕਰਨ ਦੀ ਲੋੜ ਨਹੀਂ ਹੈ ਕਿ ਉਹ ਤੁਹਾਡੀ ਮਸ਼ੀਨ ਦੀ ਮੈਮੋਰੀ ਵਿੱਚ ਆ ਜਾਣਗੇ।

ਪਲੇਟਫਾਰਮ ਅਤੇ API ਲਚਕਤਾ

TensorSharp Windows, macOS, ਅਤੇ Linux 'ਤੇ ਚੱਲਦਾ ਹੈ। ਨਵਾਂ Vulkan backend ਮੌਜੂਦਾ CUDA ਅਤੇ Metal paths ਦੇ ਨਾਲ ਉਸ ਮੈਟ੍ਰਿਕਸ ਵਿੱਚ ਬੜੀ ਸਫਾਈ ਨਾਲ ਫਿੱਟ ਹੋ ਜਾਂਦਾ ਹੈ। ਇੰਜਣ OpenAI ਅਤੇ Ollama APIs ਦੋਵਾਂ ਨਾਲ ਅਨੁਕੂਲਤਾ (compatibility) ਵੀ ਪ੍ਰਦਾਨ ਕਰਦਾ ਹੈ। ਇਹ ਚੋਣ ਇੰਟੀਗ੍ਰੇਸ਼ਨ ਦੀ ਮੁਸ਼ਕਲ ਨੂੰ ਖਤਮ ਕਰਦੀ ਹੈ। ਤੁਸੀਂ prompt templates ਨੂੰ ਦੁਬਾਰਾ ਲਿਖੇ ਬਿਨਾਂ ਜਾਂ ਨਵੇਂ response shape ਨੂੰ ਪਾਰਸ ਕੀਤੇ ਬਿਨਾਂ ਮੌਜੂਦਾ client code ਨੂੰ ਇੱਕ local TensorSharp server ਵੱਲ ਮੋੜ ਸਕਦੇ ਹੋ। ਉਹ ਟੀਮਾਂ ਜੋ ਪਹਿਲਾਂ ਹੀ ਅੰਦਰੂਨੀ ਤੌਰ 'ਤੇ Ollama ਚਲਾ ਰਹੀਆਂ ਹਨ ਜਾਂ OpenAI ਦੇ REST surface ਦੇ ਵਿਰੁੱਧ ਬਣਾ ਰਹੀਆਂ ਹਨ, ਉਨ੍ਹਾਂ ਲਈ local TensorSharp instance 'ਤੇ ਜਾਣਾ ਮੁੱਖ ਤੌਰ 'ਤੇ ਇੱਕ base URL ਬਦਲਣ ਦੀ ਗੱਲ ਹੈ।

ਉਧਾਰ ਲਈਏ ਗਏ ਅਪਟੀਮਾਈਜ਼ੇਸ਼ਨ ਜੋ ਕੰਮ ਕਰਦੇ ਹਨ

ਪ੍ਰਦਰਸ਼ਨ (Performance) ਸਿਰਫ਼ ਇਸ ਬਾਰੇ ਨਹੀਂ ਹੈ ਕਿ ਕਿਹੜਾ API GPU ਨਾਲ ਗੱਲ ਕਰਦਾ ਹੈ। TensorSharp ਕਈ ਅਜਿਹੇ ਅਪਟੀਮਾਈਜ਼ੇਸ਼ਨਾਂ ਨੂੰ ਜੋੜਦਾ ਹੈ ਜੋ ਹੋਰ ਥਾਵਾਂ 'ਤੇ production ਵਿੱਚ ਸਾਬਤ ਹੋ ਚੁੱਕੇ ਹਨ।

vLLM ਤੋਂ ਉਧਾਰ ਲਿਆ ਗਿਆ Paged KV cache, ਲੰਬੀਆਂ ਗੱਲਬਾਤਾਂ ਦੌਰਾਨ ਮੈਮੋਰੀ ਨੂੰ ਬਹੁਤ ਜ਼ਿਆਦਾ ਵਧਣ ਤੋਂ ਰੋਕਦਾ ਹੈ। ਹਰੇਕ sequence ਲਈ ਇੱਕ ਲਗਾਤਾਰ (contiguous) scratchpad ਰਾਖਵਾਂ ਰੱਖਣ ਦੀ ਬਜਾਏ, ਇੰਜਣ ਨਿਸ਼ਚਿਤ-ਆਕਾਰ ਦੇ ਪੇਜ (fixed-size pages) ਅਲਾਟ ਕਰਦਾ ਹੈ ਅਤੇ ਉਹਨਾਂ ਨੂੰ ਮੰਗ ਅਨੁਸਾਰ ਮੈ

Continuous batching, also from vLLM, improves throughput. The engine can slip new requests into active batches instead of waiting for the current group to finish. If one user’s prompt is ten tokens and another’s is two hundred, the hardware stays busier and average latency drops.

For Mixture-of-Experts models, TensorSharp implements an SSD-based cache strategy drawn from oMLX. Frequently accessed expert weights sit ready on fast storage rather than fighting for system RAM. On machines with limited memory but decent NVMe drives, this keeps MoE architectures usable.

Quantization follows the GGUF standard established by llama.cpp. Your quantized 4-bit and 5-bit models load directly without a conversion step.

The Real Takeaway

Vulkan support turns TensorSharp from an interesting C# experiment into a practical inference option for heterogeneous hardware. The roadmap is clear: validate across AMD and Intel discrete silicon, then tighten the implementation with a native Vulkan backend. If you have an AMD card in your workstation or laptop, run the build and share your results. That feedback loop is what hardens experimental code into something you can ship.

You can find the release details on the developer’s write-up. If the project saves you from juggling CUDA toolkits or wrestling with macOS version locks, leave a star on the repository. For ongoing discussion and community testing threads, the Telegram group stays open.