SMS Blog

Platform Engineering vs. DevOps Hype: Why Infrastructure Expertise Still Matters in the Cloud

DevOps changed how software is delivered. It brought the people who write code closer to the people who keep it running, and release cycles that once took months shrank to days.

But as workloads moved to the cloud, something else moved with them. Responsibility for the networks and identity systems underneath applications followed the code onto development teams, usually without much discussion.

Few organizations ever made a deliberate choice about who should design the foundations of their cloud environment. So who ended up doing it?


Where “You Build It, You Run It” Came From

In 2006, Amazon CTO Werner Vogels spoke with computer scientist Jim Gray for an interview in ACM Queue, describing how Amazon gave developers operational responsibility for the services they wrote.

Vogels said this had improved service quality for customers and for the technology itself. He compared it with the traditional model, where developers hand software to operations and move on. His summary became one of the most quoted lines in software: “You build it, you run it.”

The reasoning holds up; developers who support their own code in production see how it behaves under real load, and they hear directly from the people using it. That feedback makes each release better than the last.

Why the Idea Spread

The principle named a frustration many teams already felt. Handoffs between platform development and operations slowed releases, and accountability got blurry when something broke.

DevOps grew around this idea of shared ownership. Teams built continuous delivery pipelines and put internal developers on call for their own services. Atlassian credits the approach with code designed for production from the start and quicker response times when incidents occur.

Over time, the phrase took on a broader meaning. In many organizations, it came to mean that each development team should manage its own infrastructure, with no separate team responsible for it. But this misses Vogels’ larger point, which was organizational: shortening the feedback loop between builders and customers.

The interview also predates the cloud environments most organizations run today. Once the phrase traveled into public cloud, “run it” picked up a much larger scope.


How Infrastructure Landed on Development Teams

Public cloud turned servers and networks into API calls. Infrastructure as Code (IaC) tools let teams define an entire environment in text files and deploy it through the same pipelines as their applications.

That convenience changed who did the work. The reasoning went like this: if infrastructure is code, the people who write code can write infrastructure too. In many organizations, infrastructure management folded into development teams under a DevOps banner.

Writing the code was the easy part; a Terraform file that defines a subnet still needs someone who knows how large that subnet should be and what it should be allowed to reach. Those answers come from network and systems engineering, and the file doesn’t supply them.

The Responsibilities That Came With It

In a typical cloud environment, “running it” came to include work like this:

  • IP address planning: The CIDR ranges a team picks early decide whether networks can be connected or expanded later. Overlapping ranges are painful to untangle once workloads are live.
  • Routing and hybrid connectivity: Linking cloud networks to data centers and other clouds takes working knowledge of routing protocols and failover behavior.
  • DNS and load balancing: A small error here can take down every service that depends on name resolution.
  • Network segmentation: Deciding which systems are integrated is a core security and compliance control. Under deadline pressure, it’s easy to leave rules more open than intended.
  • IPv6 readiness: Dual-stack networks, which run IPv4 and IPv6 side by side, bring their own addressing and security considerations.
  • Identity and access management: Every role and policy is a potential path into the environment. Least privilege access takes deliberate design.
  • Encryption and key management: Decisions about who owns keys and how often they rotate shape security posture and audit outcomes.
  • Account and landing zone structure: A landing zone is the baseline account setup an environment is built on. How accounts or subscriptions are organized sets the boundaries for access and blast radius (how far a failure or breach can spread).

Each of these is a specialty. Network engineers and systems architects spend years building judgment in them, while software engineers spend those same years mastering their own craft. Expecting one team to carry both at depth is a tall order.

What It Costs Developer Time

The workload shows up in how developers spend their weeks. IDC research found that application development accounted for just 16% of developers’ time in 2024.

Operational tasks filled most of the remainder:

  • Infrastructure monitoring and management alone took 11% of developers’ time.
  • Implementing CI/CD processes ranked among the largest time commitments.
  • Time spent on security rose from 8% to 13% in a single year.

That’s a large share of skilled engineering capacity spent outside the work developers were hired to do. It also means infrastructure decisions often get made in the margins of a sprint, by people balancing feature deadlines at the same time.

Access controls are a simple yet critical component of infrastructure security: Zero Trust in Practice: Real Identity and Access Controls


The Instability That Followed

When infrastructure decisions are made without deep infrastructure experience, the effects show up in availability and security.

Outages Traced Back to Configuration and Change

The Uptime Institute has tracked outage causes for years, and its findings point to configuration as a recurring theme:

  • Network-related issues were the largest single cause of IT service outages.
  • Four in five respondents said better management and configuration practices could have prevented their most recent serious outage.
  • Outages from IT and networking issues rose to 23% of impactful outages in 2024. Uptime linked the rise to more complex IT and network environments, which lead to change management problems and misconfigurations.
  • In 2026, configuration and change management failures remain the most common driver of network-related outages.

These failures usually start with a routine change: someone edits a rule or adds a route without a full view of how the network fits together, and the impact spreads further than expected.

The Customer’s Half of Shared Responsibility

Cloud providers secure the infrastructure their services run on, and customers secure what they deploy and configure on top of it. The shared responsibility model is defined by security “of” the cloud and security “in” the cloud. It places operating system patching and security group firewall configuration on the customer’s side.

Identity is where gaps show most clearly. Datadog’s 2025 State of Cloud Security research found:

  • 39% of organizations still use IAM users in some capacity, and one in five relies on them exclusively.
  • 59% of AWS IAM users have an active access key older than one year.

These numbers are striking because long-lived keys never expire. Old keys tend to pile up when teams need access quickly and nobody owns the identity architecture. Short-lived, federated credentials take more design work up front, and that work rarely fits inside a feature sprint.

Built-in cybersecurity is critical. Don’t forget compliance: Engineering Compliance in Cloud Infrastructure


Why Platform Engineering Emerged

As the strain on development teams became harder to ignore, the industry looked for a way to keep the speed of DevOps while taking some of the infrastructure load off those teams. Platform engineering grew out of that search.

A Response to Cognitive Load

Cognitive load is the mental effort a task demands. When one team has to think about application logic, network design, and access policies all at once, each of those areas gets less attention than it needs.

The CNCF Platforms White Paper names reducing cognitive load on product teams as a core reason organizations build internal platforms. It also points to the value of sharing tools and knowledge across many teams, so each one isn’t solving the same problems separately.

The CNCF describes platform engineering as a way to scale DevOps principles through a unified platform that serves the whole organization. Platform engineering builds on DevOps; it keeps the ownership model and adds a dedicated team responsible for the shared foundation underneath.

In practice, an internal platform usually gives platform development teams:

  • Self-service infrastructure: Teams request approved environments and resources without waiting in a ticket queue.
  • Golden paths: These are preapproved, well-supported routes for common tasks, such as deploying a new service. Network and security settings come preconfigured.
  • Shared tooling: Deployment pipelines and monitoring work the same way across teams.
  • A team that owns the platform: Platform engineers maintain and improve the platform the way a product team maintains a product.

Adoption Has Gone Mainstream

Gartner predicted that by 2026, 80% of large software engineering organizations would establish platform engineering teams to provide reusable services and tools for application delivery.

That deadline is now here. For most organizations, the more useful question today is what their platform team needs to be capable of.


Platform Engineering as an Infrastructure Discipline

A platform is only as sound as the infrastructure decisions inside it, which puts specific kinds of expertise at the center of the discipline.

Networking at the Foundation

Network design choices get made early, and many are difficult to undo.

  • Address ranges are permanent: In AWS, you can’t remove the primary IPv4 CIDR block of a virtual private cloud (VPC).
  • Overlaps block connections: Adding address ranges to a VPC with an active peering connection requires that they don’t overlap with the peer network.

An address plan drawn up in a project’s first sprint can therefore limit how networks connect years later.

Platform engineers with enterprise networking backgrounds design for that long view. Their work typically covers:

  • IP address management that spans every account and region, so ranges never collide.
  • Routing and hybrid connectivity designed with failover in mind from the start.
  • Segmentation that reflects how data and workloads should be isolated from each other.
  • IPv6 planning that happens before it becomes urgent.

System Architecture Shapes Resilience

A landing zone is a multi-account environment that serves as the starting point for workloads. Building one involves decisions about account structure, networking, security and access management.

These decisions shape how risk is contained. AWS treats each account as a unit of security protection, so a threat in one account stays isolated from the others. Some early choices carry long-term consequences;  for example, moving an AWS Control Tower landing zone to a different home Region requires decommissioning it and working with AWS Support.

Systems architects think in failure domains, meaning the set of components that break together when something goes wrong. Designing those boundaries well limits how far an outage or breach can spread.

Infrastructure Code Written by Infrastructure Engineers

Infrastructure as Code records decisions. The quality of the environment depends on the quality of those decisions.

Platform team members turn infrastructure expertise into reusable modules. When a developer requests a network, they receive one that’s already sized and secured correctly. The DevOps practices that make this work include:

  • Versioned modules for common components, maintained by people who understand what each one does.
  • Policy as code, meaning automated rules that check every change before it deploys.
  • State management, which keeps an accurate record of what the IaC tool has deployed.
  • Drift detection, which flags when the live environment no longer matches its code.

Default Security Hardening

When security guardrails live in the platform, every team inherits them automatically. Common platform-level controls include:

  • Short-lived, federated credentials for both people and pipelines, which replace static keys.
  • Least-privilege roles generated from reviewed templates.
  • Organization-wide policies that block risky actions before they happen. AWS service control policies are most often used to protect shared infrastructure, landing zones, and core security services.
  • Logging and detection enabled in every account from the day it’s created.

Security teams can then review the platform’s design once and have those controls apply everywhere it’s used.

What Development Teams Gain

Developers keep full ownership of their applications; they still deploy, operate and support what they build, so the feedback loop Vogels described stays intact.

The difference is the ground they’re standing on. The networks, accounts and access policies underneath their services come from people who specialize in them.


Questions to Ask About Your Environment

A short internal review can show where infrastructure ownership sits today. These questions are a good place to start:

  • Who owns IP address management across every account and region? If the answer varies by team, overlapping ranges are likely already present or on their way.
  • Who reviews IaC changes for network and security impact before they merge? Code review that checks syntax alone can miss the decisions that matter most.
  • How quickly would you know about a manual change in production? Drift that goes unnoticed for weeks tends to surface during the next incident.
  • Where do your pipelines get their cloud credentials, and how long do those credentials last? Static keys in CI/CD systems are a common entry point for attackers.
  • Does your platform team include engineers with networking and systems architecture backgrounds? Tooling skills and infrastructure design skills are different specialties.
  • If your landing zone needed a redesign, who would lead it? An unclear answer usually means nobody owns the foundation.

Uncertain answers are common, and they’re very useful. Each one points to an infrastructure decision currently being made by default.


Give Infrastructure an Expert Owner

Cloud providers will keep releasing new services and abstracting away more of the stack. Every one of those services still runs on addressing, routing and identity decisions that someone has to make.

Organizations that give those decisions a dedicated, qualified owner put their development teams on steadier ground. That shift starts with recognizing infrastructure as its own discipline and staffing it that way.

Long before hyperscalers existed, SMS was building data centers and enterprise networks. That background includes deep work in IPv4 and IPv6, network service provider technologies, and a long partnership with Cisco (ask us about this one). Our platform engineering team brings the same expertise to commercial organizations running in the cloud.

If you’re unsure who owns the foundations of your cloud environment, talk to our platform engineers. We’ll help you get the infrastructure right so your development teams can build on it.

Picture of Andrew Stanley

Andrew Stanley

Andrew Stanley, SMS' Chief Technology Officer, joined in 2002 as a junior network engineer, supporting Department of Defense IT infrastructures and leading programs for the Executive Office of the President and DARPA. Promoted to Director of Engineering in 2021, he drove talent development and innovation across the company. A private pilot at 16 and former U.S. Army Information Systems Analyst, Andrew earned an IT degree from George Mason University through the Army's Green to Gold program. View Andrew's LinkedIn

Leave a Reply

Your email address will not be published. Required fields are marked *