AI & Automation
Devix AI Runtime
A from-scratch C++ GPT training and inference runtime with CPU/CUDA/DirectML backends, its own .devix model format, and an OpenAI-style API.
- Client
- Devix — intellectual property (Hussain AlZadjali)
- Completed
- April 2025
- Technologies
- C++17CUDADirectMLOpenBLASFAISSNext.jsGradio
Devix AI Runtime is a proprietary GPT runtime written from scratch in C++ with CPU, CUDA, and DirectML backends — giving inference sovereignty and AMD GPU compatibility beyond NVIDIA-only stacks. It defines its own .devix model format, exposes an OpenAI-style REST API, ships a Devix AI Studio, and includes a validated government tender/RFP analyzer POC. All 36 runtime tests pass on AMD GPUs.
36/36 pass
GPU tests
CPU/CUDA/DirectML
Backends
C++17
Language
OpenAI-style
API
The challenge
Devix needed inference sovereignty and hardware flexibility — running its own models on AMD GPUs and CPUs without depending on NVIDIA-only or cloud inference stacks.
The solution
A hand-built C++17 runtime with pluggable CPU/CUDA/DirectML backends, a compact .devix model format, an OpenAI-compatible REST API for drop-in integration, a Studio UI, and an integrated FAISS-backed tender RAG analyzer proving the runtime on a real government use case.
Architecture
C++17 core with OpenBLAS (CPU), CUDA, and DirectML (AMD) backends; .devix model serialization; OpenAI-style REST server; FAISS retrieval for the RAG POC; and a Gradio/Next.js studio shell.
Key features
- From-scratch C++17 GPT runtime
- CPU, CUDA, and DirectML (AMD) backends
- Custom .devix model format
- OpenAI-compatible REST API
- Devix AI Studio interface
- Validated tender/RFP analyzer POC
Have a similar problem to solve?
Tell us what you're building — we'll tell you honestly whether we're the right team for it.
Start a conversation