All work

AI & Automation

Devix AI Runtime

A from-scratch C++ GPT training and inference runtime with CPU/CUDA/DirectML backends, its own .devix model format, and an OpenAI-style API.

Client
Devix — intellectual property (Hussain AlZadjali)
Completed
April 2025
Technologies
C++17CUDADirectMLOpenBLASFAISSNext.jsGradio

Devix AI Runtime is a proprietary GPT runtime written from scratch in C++ with CPU, CUDA, and DirectML backends — giving inference sovereignty and AMD GPU compatibility beyond NVIDIA-only stacks. It defines its own .devix model format, exposes an OpenAI-style REST API, ships a Devix AI Studio, and includes a validated government tender/RFP analyzer POC. All 36 runtime tests pass on AMD GPUs.

36/36 pass

GPU tests

CPU/CUDA/DirectML

Backends

C++17

Language

OpenAI-style

API

The challenge

Devix needed inference sovereignty and hardware flexibility — running its own models on AMD GPUs and CPUs without depending on NVIDIA-only or cloud inference stacks.

The solution

A hand-built C++17 runtime with pluggable CPU/CUDA/DirectML backends, a compact .devix model format, an OpenAI-compatible REST API for drop-in integration, a Studio UI, and an integrated FAISS-backed tender RAG analyzer proving the runtime on a real government use case.

Architecture

C++17 core with OpenBLAS (CPU), CUDA, and DirectML (AMD) backends; .devix model serialization; OpenAI-style REST server; FAISS retrieval for the RAG POC; and a Gradio/Next.js studio shell.

Key features

  • From-scratch C++17 GPT runtime
  • CPU, CUDA, and DirectML (AMD) backends
  • Custom .devix model format
  • OpenAI-compatible REST API
  • Devix AI Studio interface
  • Validated tender/RFP analyzer POC

Have a similar problem to solve?

Tell us what you're building — we'll tell you honestly whether we're the right team for it.

Start a conversation
Devix AI Runtime | C++ GPT Engine | Devix — Devix