Skip to content
shamimnael
Shamim Nael

Shamim Nael · Berlin, Germany

About

I never left the network. I built upwards from it.

Two decades spanning carrier networks, platform engineering, and the infrastructure that machine learning now depends on. The three turned out to be one discipline.

I spent the first half of my career designing and operating carrier networks: the backbones that national mobile and broadband operators run on. Work at that scale teaches a particular discipline. You think in failure domains. You understand the blast radius of a change before you make it. You assume that every decision will eventually be examined by someone who was not in the room.

When I moved into platform and site-reliability work, I expected most of that to be irrelevant. The opposite proved true. Platform teams meet capacity planning, multi-tenancy, graceful degradation and change control under load as though they were new problems. Networks resolved them decades ago, and the reasoning transfers almost intact.

Today I work on infrastructure where the machine-learning platform sits on a stack I also operate: the network beneath it, the identity layer, the data services, the monitoring. Not because managed services are the wrong choice, since often they are the right one, but because understanding a system all the way down is the difference between operating it and hoping it holds.

That is the subject of this site. What it takes to run AI systems on infrastructure you control, and what happens when those same systems are applied back to the operational work underneath. Written from practice, with the reasoning shown and the limitations stated.

Experience

  1. 2024–present

    Yoummday

    DevOps Engineer

    Design and operation of a self-hosted production estate: networking and security, highly available data services, real-time communications, the observability stack, and the machine-learning platform running on top of it. Defined as code, end to end.

  2. 2022–2024

    Agileful, then Cocus

    Senior DevOps / Systems Engineer

    Cloud infrastructure and delivery on AWS: infrastructure as code, container orchestration, continuous deployment, and production on-call. The point at which network engineering and platform engineering stopped being separate disciplines for me.

  3. 2015–2022

    Rayaneh Tahlilgaran Ozhan

    Network Operations Engineering Manager

    Led the network operations team serving enterprise and carrier customers. Software-defined WAN deployment, data-center interconnect, IPv6 transition programs, and the first systematic automation of configuration management and compliance checking.

  4. 2012–2015

    Mobinnet Telecom

    Senior Network Design & Operation Engineer

    Design and operation of the IP/MPLS network for a national fixed-wireless broadband provider. End-to-end quality of service, layer-3 VPN services for enterprise customers, and the network redesign that supported the move to LTE.

  5. 2007–2012

    Nokia / Huawei

    IP Backbone Engineer & Team Lead

    Design and second-line support for a national mobile operator’s IP backbone. MPLS VPN and traffic engineering across the core, integration of mobile core elements into the transport network, and quality-of-service policy for converged traffic.

Credentials

  • CCIE ×2
  • MSc Computer Science with Artificial Intelligence
  • Two decades in infrastructure
  • Berlin, Germany

Follow the writing

New posts, over RSS.

New posts go out over RSS. If you would rather talk than read, the inbox is open.