This project empirically demonstrates the adverse effects of CPU limits in Kubernetes by analyzing cgroup throttling statistics when Pods are subjected to heavy, sustained load.
We deploy an identical, CPU-intensive server application into a Kind cluster (Kubernetes 1.35) across four distinct Deployment configurations and use Fortio to flood them with requests, observing the resulting throttling via Prometheus and Grafana.
- π Project Goal
- π¦ Directory Structure
- π§ͺ The Four Server Deployment Cases
- π οΈ Setup and Execution
- π Analysis and Observation
- ποΈ Cleanup
- π» Linux CPU Scheduler and Time Slices
- βWhen to use CPU limits?
This project is designed to use specific cgroup metrics to prove why setting CPU limits is detrimental to application performance and should generally be avoided, aligning with best practices that suggest allowing Pods to burst when extra node capacity is available.
We will visualize and quantify the performance characteristics and resource consumption using highly specific resource utilization data and cgroup throttling statistics collected directly from the server Pods.
Here's a visual representation of the problem CPU limits can cause:
The four key resource configurations being tested are:
- No Limits or Requests: (The worst case for stability)
- Request Only: (The recommended configuration for maximum burst performance)
- Limit Only: (Hard-capped performance)
- Request equals Limit: (Hard-capped, but with resource guarantee)
To precisely demonstrate CPU throttling, the Go server application exposes specific resource utilization and cgroup metrics in its response logs. These metrics are crucial for quantifying the difference between available CPU time and actual CPU time consumed (due to throttling).
| Metric Key | Description | Relevance to Project Goal |
|---|---|---|
rusage_user_seconds |
The amount of time spent by the process in user mode (CPU time). | Measures actual work done by the application logic. |
rusage_system_seconds |
The amount of time spent by the process in kernel mode. | Measures time spent by the OS on behalf of the application (e.g., I/O). |
num_of_cores |
Total CPU consumption in cores (calculated as (user + system) / elapsed_seconds). |
Shows the average CPU cores consumed by the process, directly revealing throttling effects. |
cgroup.nr_periods_delta |
Total number of scheduling periods elapsed. | Baseline for cgroup activity. |
cgroup.nr_throttled_delta |
The number of times the process was throttled (i.e., paused due to hitting a CPU limit). | CRITICAL: Direct count of throttling incidents. Expected to be high for Pods with CPU limits. |
cgroup.throttled_time_ns_delta |
Cumulative time (in nanoseconds) the process has been throttled. | CRITICAL: The total time the application was suspended due to the CPU limit. |
cgroup.cpu_limit |
The detected CPU limit value, expressed in cores (if quota/period are set). | Confirms the limit applied by Kubernetes/cgroups. |
| Directory | Purpose | Key Files |
|---|---|---|
server/ |
The Go application that simulates a CPU-intensive workload. | Dockerfile, main.go |
deploy/ |
Kubernetes manifests configured via Kustomize: four server Deployments, per-server Services, Fortio load generator, Grafana dashboard, and kube-prometheus-stack via Helm. | kustomization.yaml, server.yaml, fortio.yaml, grafana-dashboard.yaml, kind-config.yaml |
The deploy/server.yaml manifest creates four Deployments, each with a dedicated Service, to test different resource configurations:
| Case | Deployment | requests.cpu |
limits.cpu |
QoS Class | Expected Behavior Under Stress |
|---|---|---|---|---|---|
| 1. No Requests/Limits | server-1 |
β | β | BestEffort | Unpredictable performance; first to be evicted under resource pressure. Can consume all available CPU. |
| 2. Request Only | server-2 |
3 | β | Burstable | Receives guaranteed baseline CPU share. Can burst beyond the request to use all available CPU capacity on the node. |
| 3. Limit Only | server-3 |
β | 3 | Guaranteed* | No explicit request, but hard-capped at the limit value. Will experience throttling immediately upon hitting the cap. |
| 4. Request + Limit | server-4 |
1 | 3 | Burstable | Receives a lower CPU share, but is hard-capped at the limit and will be throttled when it attempts to use more. |
* Setting only limits implicitly sets requests = limits.
You must have the following tools installed and available in your PATH:
- Docker (or compatible container runtime)
- kind >= v0.31.0 (Kubernetes IN Docker)
- kubectl (Kubernetes command-line tool)
- kustomize (Kubernetes manifest composition)
- helm (Kubernetes package manager)
- make (or run the commands manually)
-
Clone the Repository:
git clone https://github.com/michalschott/kubernetes-resources-cpu.git cd kubernetes-resources-cpu -
Build, Deploy, and Run: Use the provided
Makefileto handle all steps: tool checks, image builds, Kind cluster creation (Kubernetes 1.35), metrics server deployment, Prometheus + Grafana stack, and application deployment.make deploy
-
Verify Cluster Status: Ensure all four server Deployments and Fortio are ready.
kubectl get pods -n stresstest
-
Open Grafana: Port-forward Grafana and open the pre-provisioned "CPU Throttling Demo" dashboard.
make port-forward
- URL: http://localhost:3000
- Credentials:
admin/prom-operator - Navigate to: Dashboards β CPU Throttling Demo
-
Run the Load Test: In a separate terminal, blast all four servers with Fortio (10 min, 10 QPS, 5 connections per server):
make load-start
For a shorter test (30s):
make load-quick
-
Swap Server Configurations: Switch between the two server manifests while the load test is running (zero-downtime rolling updates):
make swap-default # Mixed resources: no res, request only, limit only, request+limit make swap-fair # All equal: requests.cpu=1, no limits
Once the load test is running, use the CPU Throttling Demo Grafana dashboard to observe:
-
CPU Usage (cores): Pods without limits (
server-1,server-2) burst freely. Pods with limits (server-3,server-4) flatline at their quota. -
Throttle Ratio:
server-3andserver-4approach 70-80% throttled periods.server-1andserver-2show 0%. -
CPU Usage vs Request: Shows actual usage alongside request and limit lines β visually demonstrates bursting above requests and capping at limits.
-
Monitor with kubectl top:
kubectl top pods -n stresstest kubectl top node
-
Live Resource Changes: Since servers are Deployments, you can modify resources on the fly:
# Add a limit to the previously unlimited server-2 kubectl -n stresstest set resources deployment/server-2 --limits=cpu=3 # Remove the limit from server-3 kubectl -n stresstest set resources deployment/server-3 --limits=cpu=0 --requests=cpu=3
Watch the Grafana dashboard as the rolling update takes effect β throttling appears or disappears within seconds.
-
Scale down to observe redistribution:
# Remove the bursting server-2 kubectl -n stresstest scale deployment/server-2 --replicas=0Observe the freed CPU capacity on the dashboard.
server-1(also unlimited) immediately absorbs the released capacity.# Remove server-1 as well kubectl -n stresstest scale deployment/server-1 --replicas=0The remaining limited Pods (
server-3andserver-4) cannot use the freed capacity due to their hard quotas β the CPU sits idle and wasted.π‘ Conclusion: CPU capacity released by bursting Pods is redistributed among other bursting Pods. But when only limited Pods remain, freed capacity goes unused β resulting in wasted, paid-for compute resources.
To remove the Kubernetes cluster and clean up all deployed resources:
make cleanThe core mechanism governing how container workloads receive CPU time is the Linux kernel's Completely Fair Scheduler (CFS). The CFS is an implementation of a proportional share scheduler, meaning it aims to give every runnable task (thread) a "fair" and equal share of the CPU time, unless weights are applied. This fairness is achieved by allocating CPU time in short intervals called time slices or quantums.
When a task is ready to run, the CFS places it into a red-black tree (a data structure used for efficient scheduling) and tracks its accumulated runtime, or vruntime (virtual runtime). The scheduler constantly picks the task with the lowest vruntime to run next. When a task runs for its allocated time slice, its vruntime increases, pushing it further down the priority list and allowing another task to run. This continuous, rapid switching gives the appearance of simultaneous execution.
In the context of Kubernetes CPU Requests, the proportional share is modified: a higher CPU request translates into a greater scheduling weight (cpu.shares or cpu.weight), which effectively grants the container more frequent or longer time slices, ensuring it receives its guaranteed share of CPU time, especially under contention.
This time-slice mechanism ensures that setting a CPU Limit for a container means its total CPU consumption (the sum of its time slices) within a given period (e.g., 100ms) can never exceed its hard-capped quota. Once the quota is used up, the container is forcibly suspended until the next period begins, which is precisely what is measured by your cgroup.throttled_time_ns_delta metric.
To explain how CPU Limits cause the throttling, we need to focus on the Linux kernel mechanism called CFS Bandwidth Control. This is a separate cgroup feature from the proportional sharing used for CPU Requests.
A Kubernetes CPU Limit is enforced by the kernel using two specific parameters within the container's control group:
cpu.cfs_period_us(CFS Period): This is a fixed time window, usually set to 100,000 microseconds (100ms). This is the accounting window for CPU usage.cpu.cfs_quota_us(CFS Quota): This is the total amount of CPU time (in microseconds) the container's processes are allowed to consume within the defined period.
The Kubelet (via the container runtime) translates your Kubernetes CPU Limit into this quota using a simple formula:
Quota = CPU Limit (in cores) Γ Period
Example: If you set a CPU Limit of 500m (0.5 CPU), the quota is calculated as:
Quota = 0.5 Γ 100ms = 50ms
This quota creates a hard cap on the container's CPU usage, regardless of available capacity on the Node:
- As soon as the container's processes run, they consume the 50ms quota for the current 100ms period.
- If the container attempts to use more than 50ms of CPU time before the 100ms period is over, the kernel's scheduler will immediately suspend (throttle) all of its threads.
- The threads are prevented from running, even if the Node has many idle CPU cores. They remain suspended until the start of the next 100ms period, at which point the quota is refilled, and the processes are allowed to run again.
This mechanism explains why Request Only pod (server-2) can burst and never throttles (since no quota is set), while Limit Only (server-3) and Request lower than Limit (server-4) pods will throttle the instant their quota is exhausted under heavy load, severely degrading their performance.
β Setting Kubernetes CPU limits is generally not considered a best practice because they are enforced using the Linux kernel's CFS Bandwidth Control, which can lead to severe performance degradation through CPU throttling.
When a container hits its hard limit, it is forcibly suspended (throttled) until the next time period, even if the node has abundant idle CPU capacity. This artificial suspension significantly increases latency and reduces application throughput, which is measured by metrics like cgroup.throttled_time_ns_delta.
β The recommended approach is to set CPU requests only, which guarantees a proportional share of CPU time during contention while allowing the workload to burst and utilize any available spare capacity on the node, maximizing resource efficiency.
| Argument | Why It Doesn't Work | Better Alternative |
|---|---|---|
| Hard Multi-Tenancy/Isolation | CFS weights from requests already guarantee each pod its proportional share of CPU under contention. Limits waste idle capacity that other tenants could use. |
Dedicated node pools with taints/tolerations. |
| Cost Control/Optimization | Limits do not reduce costs β node count and cloud spend are driven by the sum of requests, not limits. Limits actually increase costs: pods that cannot burst require more provisioned capacity to handle the same throughput. |
Right-size requests and use Karpenter for optimal node selection. |
| Debugging Runaway Processes | In practice, nobody surgically applies a CPU limit during an incident. | Restart or scale down the misbehaving pod and fix the root cause. |
There is no practical scenario in a well-managed Kubernetes environment where CPU limits improve outcomes. They only add artificial latency and waste paid-for compute capacity.
