Skip to content
shamimnael

Infrastructure engineer · Berlin, Germany

From carrier backbones to self-hosted AI.

I design and operate the infrastructure that AI systems depend on: networks, platforms, and the automation that holds them together. I write about running those systems on infrastructure you control.

Large-scale network design

Architecture for carrier and enterprise networks. Hierarchy and modularity that keep failure domains contained, routing designs that converge predictably under load, and capacity models that hold up against growth nobody forecast. Multi-vendor by necessity, and designed to be operated rather than only deployed.

Platform engineering & automation

Infrastructure as code across cloud and self-hosted estates. The discipline is less about tooling than about making the state of an estate knowable: one authoritative source of inventory, changes that can be reviewed before they land, and drift that surfaces before it becomes an incident.

AI platform operations

The systems machine learning depends on once it leaves a notebook: experiment tracking, vector search, workflow orchestration and model observability, over the virtualization, identity and monitoring layers beneath them. Assembling the components is rarely the difficulty; keeping them dependable together is.

  • CCIE ×2
  • MSc Computer Science with Artificial Intelligence
  • Two decades in infrastructure
  • Berlin, Germany

The work

Two directions, one discipline.

I operate the platforms that AI systems run on, and I apply those same systems back to the operational work underneath. Each direction informs the other.

Infrastructure for AI

Building the ground AI stands on

The platforms machine learning and language-model applications actually run on: experiment tracking, vector search and workflow orchestration, together with the virtualization, identity and observability layers that make them dependable enough to trust in production.

AI for infrastructure

Turning those systems back on the work

Applying language models to operational problems: configuration analysis, runbook and documentation retrieval, incident triage. Assessed honestly, including the cases where the added complexity is not repaid.

Follow the writing

New posts, over RSS.

New posts go out over RSS. If you would rather talk than read, the inbox is open.