~/srekubecraft zsh

$ neofetch

nick@srekubecraft

Nick Nikolakakis

--------------------------
role
Principal SRE, Platform & AI Engineer
focus
Kubernetes · CNCF · Platform · AI
building
AI agents · LLMOps · MCP
shell
zsh
since
2011

// The SRE Diary: Chronicles of Cloud, Kubernetes and AI

~/posts/auditing-proxies-and-meshes-for-the-kubectl-websockets-transport 21 min · 4468 words

// Kubernetes is retiring the SPDY streaming transport behind kubectl exec, attach, cp and port-forward in favour of WebSockets. This is an audit you can run against your clusters this week to find the hop that will break, before the upgrade finds it for you.

~/posts/kubernetes-gang-scheduling-workload-api 17 min · 3485 words

// KEP-4671 landed gang scheduling in Kubernetes 1.37 beta, behind a feature gate that is off by default. A four-node Kind cluster with six fake DRA GPUs shows the partial-placement deadlock, the single field that fixes it, and the two limits nobody documents: the Job field is pruned silently without a second alpha gate, and workload-aware preemption cannot reclaim DRA devices.

~/posts/modelplane-beyond-placement-routing-dynamo-gotchas 22 min · 4548 words

// Part two of the Modelplane fleet inference series, covering the layer above the scheduler. ModelService is a routing policy rather than a load balancer, spec.stack Dynamo swaps the whole gang-scheduling machinery, and four documented gotchas will cost you an afternoon each.

~/posts/modelplane-fleet-inference 29 min · 5971 words

// A hands-on tour of Modelplane, the Crossplane-based control plane for AI inference across a fleet of GPU clusters. Two regions on Kind with fake DRA GPUs, the platform/developer contract written in DRA's own vocabulary, fleet-level drain, and the release trap that costs you two APIs.

~/posts/llm-d-distributed-inference 21 min · 4271 words

// A hands-on tour of llm-d, the CNCF Sandbox framework for distributed LLM inference on Kubernetes - inference-aware routing, prefill/decode disaggregation, and KV-cache offload. Includes a GPU-free demo on Kind using the vLLM simulator, wired with Flux GitOps.

~/posts/choragos-multi-agent-orchestrator 26 min · 5350 words

// Choragos runs a team of AI coding agents with a real division of labour: an orchestrator that plans and delegates, workers with their own context, model and credentials, and a delegate/work-done protocol carrying work between them. A walkthrough of v0.11.2 with config recipes for real teams.

~/posts/introducing-sphragis-eu-ai-act-compliance-gateway 17 min · 3410 words

// Sphragis is a self-hosted Go gateway that strips PII out of every LLM request and response before it leaves your network and writes a tamper-evident, hash-chained audit log. A walkthrough of the v0.3.0 release: local redaction, reversible tokenization, multi-provider routing, and OpenTimestamps anchoring.