[{"data":1,"prerenderedAt":329},["ShallowReactive",2],{"blog-en-kubernetes-cpu-throttling-connection-pool-exhaustion":3,"blog-related-en-kubernetes-cpu-throttling-connection-pool-exhaustion":277,"blog-en-kubernetes-cpu-throttling-connection-pool-exhaustion-alt":265},{"id":4,"title":5,"author":6,"body":7,"date":259,"description":260,"extension":261,"image":262,"locale":263,"meta":264,"navigation":265,"path":266,"seo":267,"stem":268,"tags":269,"__hash__":276},"blog\u002Fblog\u002Fen\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion.md","The Trap of 'CPU Usage Is Low, So We're Fine': The Hidden Ceiling of Throttling and Connection Pool Exhaustion Kubernetes Won't Show You","Kubo Team",{"type":8,"value":9,"toc":245},"minimark",[10,15,27,30,41,61,67,76,80,86,89,103,111,123,127,133,146,155,158,162,168,173,181,185,188,192,201,205,208,214,222,226,229,232],[11,12,14],"h2",{"id":13},"cpu-usage-is-low-but-its-slow-the-classic-way-teams-misdiagnose-kubernetes-latency","\"CPU Usage Is Low, But It's Slow\" — The Classic Way Teams Misdiagnose Kubernetes Latency",[16,17,21],"p",{"className":18,"dir":20},[19],"content-paragraph","ltr",[22,23],"img",{"src":24,"alt":25,"width":26,"height":26},"https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion\u002Fsection01.webp","","inherit",[16,28,29],{},"You open the production dashboard and CPU usage sits at around 30%. Yet API response times are clearly slow. This \"contradictory dashboard\" is the first trap that leads many infrastructure engineers to misdiagnose the true cause of Kubernetes latency.",[16,31,32,33,40],{},"The root cause lies in how CPU limits are enforced at the kernel level. According to the ",[34,35,39],"a",{"href":36,"rel":37},"https:\u002F\u002Fkubernetes.io\u002Fdocs\u002Fconcepts\u002Fconfiguration\u002Fmanage-resources-containers\u002F",[38],"nofollow","official Kubernetes documentation",", CPU limits are enforced by the CFS (Completely Fair Scheduler), and once a container approaches its limit, the kernel restricts its access to the CPU itself. Unlike memory, where a process gets killed by OOM, there's no obvious error — just a vague sense that things are \"slow.\"",[16,42,43,44,48,49,54,55,60],{},"What makes this worse is that the \"CPU usage\" shown by most monitoring tools is measured ",[45,46,47],"strong",{},"after"," throttling has already occurred. As ",[34,50,53],{"href":51,"rel":52},"https:\u002F\u002Fwww.datadoghq.com\u002Fblog\u002Fkubernetes-cpu-requests-limits\u002F",[38],"Datadog's blog post"," points out, a container's usage looks low precisely because it's being restricted — the intuition that \"low usage means there's headroom\" simply doesn't hold here. ",[34,56,59],{"href":57,"rel":58},"https:\u002F\u002Fwww.ibm.com\u002Fthink\u002Ftopics\u002Fkubernetes-cpu-throttling-identification",[38],"IBM's troubleshooting guide"," similarly notes that investigations focused only on CPU usage tend to miss throttling entirely.",[16,62,64],{"className":63,"dir":20},[19],[22,65],{"src":66,"alt":25,"width":26,"height":26},"https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion\u002Fsection02.webp",[16,68,69,70,75],{},"Investigating this kind of \"latency that contradicts what the dashboard shows\" takes a long time if your monitoring stack isn't already in place. With a managed K3s environment like ",[34,71,74],{"href":72,"rel":73},"https:\u002F\u002Fkubo.hexabase.io\u002F",[38],"Kubo",", which comes with Prometheus and Grafana built in, you don't need to spend extra time building out this kind of observability from scratch.",[11,77,79],{"id":78},"why-adding-more-pods-scaling-with-hpa-can-make-things-worse","Why Adding More Pods (Scaling with HPA) Can Make Things Worse",[16,81,83],{"className":82,"dir":20},[19],[22,84],{"src":85,"alt":25,"width":26,"height":26},"https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion\u002Fsection03.webp",[16,87,88],{},"Suspecting CPU throttling, you scale out by adding more replicas — but latency doesn't improve, or even gets worse. This usually means the bottleneck actually lives in the application layer or the database layer.",[16,90,91,92,96,97,102],{},"A classic example is database connection pool exhaustion. PostgreSQL manages the number of concurrent connections via the ",[93,94,95],"code",{},"max_connections"," parameter, and as ",[34,98,101],{"href":99,"rel":100},"https:\u002F\u002Fplanetscale.com\u002Fblog\u002Fscaling-postgres-connections-with-pgbouncer",[38],"PlanetScale explains",", Postgres forks an OS process for every connection, so memory usage and context-switching costs rise sharply as connections grow. When HPA adds more Pods on top of this, and each Pod maintains its own connection pool, the total number of connection requests balloons even further.",[16,104,105,110],{},[34,106,109],{"href":107,"rel":108},"https:\u002F\u002Fcubeapm.com\u002Fblog\u002Fpostgresql-connection-pool-exhausted-kubernetes\u002F",[38],"CubeAPM's technical article"," walks through exactly this pattern of connection exhaustion in Kubernetes-based applications. Adding Pods is supposed to increase capacity — but against a database's connection limit, it backfires, triggering a surge of \"too many clients\" errors instead.",[16,112,113,114,119,120,122],{},"A well-known way to avoid this is a connection pooler. As ",[34,115,118],{"href":116,"rel":117},"https:\u002F\u002Fneon.com\u002Fdocs\u002Fconnect\u002Fconnection-pooling",[38],"Neon's official documentation"," explains, placing a pooler like PgBouncer in between lets you multiplex many application-level connections down to a small number of real database connections, allowing Pods to scale horizontally without ever hitting the ",[93,121,95],{}," ceiling.",[11,124,126],{"id":125},"the-metric-you-should-actually-be-watching-isnt-cpu-how-to-isolate-the-real-cause-of-kubernetes-latency","The Metric You Should Actually Be Watching Isn't CPU% — How to Isolate the Real Cause of Kubernetes Latency",[16,128,130],{"className":129,"dir":20},[19],[22,131],{"src":132,"alt":25,"width":26,"height":26},"https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion\u002Fsection04.webp",[16,134,135,136,141,142,145],{},"If you only ever look at CPU usage, you risk missing both throttling and connection pool exhaustion at once. ",[34,137,140],{"href":138,"rel":139},"https:\u002F\u002Flast9.io\u002Fblog\u002Fkubernetes-cpu-throttling\u002F",[38],"Last9's engineering blog"," explains that the Prometheus metric ",[93,143,144],{},"container_cpu_cfs_throttled_seconds_total"," — not CPU usage — directly shows how much time was actually spent throttled. Even when CPU usage looks low, a high value here is a red flag for throttling.",[16,147,148,149,154],{},"On the database side, isolating latency requires visibility into connection pool utilization and queue wait times. If HPA is configured to scale on CPU alone, ",[34,150,153],{"href":151,"rel":152},"https:\u002F\u002Fdocs.cloud.google.com\u002Fkubernetes-engine\u002Fdocs\u002Fconcepts\u002Fhorizontalpodautoscaler",[38],"Google Cloud's official documentation"," shows that since Kubernetes 1.6, the Custom Metrics API lets you feed application-specific metrics — like queue length or response time — into HPA's scaling conditions via Prometheus. That's the first step away from a setup that scales purely on CPU.",[16,156,157],{},"In short, correctly identifying the true cause of Kubernetes latency requires looking at a minimum of four metrics side by side: CPU%, CFS throttled time, database connection pool utilization, and p99 latency. Relying on just one of these will always leave you blind to at least one other possible cause.",[11,159,161],{"id":160},"the-right-fix-what-to-do-before-you-add-more-pods","The Right Fix: What to Do Before You Add More Pods",[16,163,165],{"className":164,"dir":20},[19],[22,166],{"src":167,"alt":25,"width":26,"height":26},"https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion\u002Fsection05.webp",[169,170,172],"h3",{"id":171},"consider-removing-cpu-limits-entirely-for-latency-sensitive-services","Consider Removing CPU Limits Entirely for Latency-Sensitive Services",[16,174,175,180],{},[34,176,179],{"href":177,"rel":178},"https:\u002F\u002Fwww.groundcover.com\u002Fblog\u002Fkubernetes-cpu-throttling",[38],"groundcover's analysis"," points out that setting CPU limits too aggressively can cause containers to be throttled unnecessarily even when the node has spare capacity. Rather than applying a CPU limit to every workload, it's worth considering securing priority via requests while relaxing or removing limits for latency-sensitive services.",[169,182,184],{"id":183},"switch-hpas-scaling-trigger-from-cpu-to-application-level-metrics","Switch HPA's Scaling Trigger from CPU to Application-Level Metrics",[16,186,187],{},"Using the Custom Metrics API described above, switching from CPU-based scaling to scaling on request queue length or response time lets you catch the \"CPU is free but things are backed up\" state much earlier.",[169,189,191],{"id":190},"put-a-connection-pooler-in-the-middle","Put a Connection Pooler in the Middle",[16,193,194,195,200],{},"As ",[34,196,199],{"href":197,"rel":198},"https:\u002F\u002Fkomodor.com\u002Flearn\u002Fkubernetes-cpu-limits-throttling\u002F",[38],"Komodor's article"," also points out, misdiagnosing resource issues consumes a huge amount of operational time. Placing a pooler like PgBouncer as a sidecar or intermediary layer decouples Pod count growth from database connection count growth.",[169,202,204],{"id":203},"size-maxreplicas-backwards-from-what-the-database-can-actually-accept","Size maxReplicas Backwards from What the Database Can Actually Accept",[16,206,207],{},"Instead of setting HPA's maxReplicas based solely on cluster CPU capacity, calculate it backwards from \"the number of connections the database can safely accept ÷ the number of connections per Pod.\" This prevents scale-out from becoming a self-inflicted source of errors.",[16,209,211],{"className":210,"dir":20},[19],[22,212],{"src":213,"alt":25,"width":26,"height":26},"https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion\u002Fsection06.webp",[16,215,216,217,221],{},"Having all of these metrics visible on a single screen from day one is the fastest way to avoid a drawn-out misdiagnosis. ",[34,218,220],{"href":72,"rel":219},[38],"Kubo Cloud"," is a managed Kubernetes service built on K3s with Prometheus and Grafana included by default, so you can start tracking CFS throttled time and setting up custom-metric-based HPA immediately, without building an observability stack from scratch.",[11,223,225],{"id":224},"conclusion-scaling-isnt-magic","Conclusion: Scaling Isn't Magic",[16,227,228],{},"\"Add more Pods\" is a convenient fix, but it isn't a cure-all. When the real bottleneck is kernel-level CPU throttling or database-side connection pool exhaustion, scaling out doesn't just fail to fix the problem — it can actively introduce a new one in the form of connection errors.",[16,230,231],{},"Correctly identifying the cause of Kubernetes latency starts with not trusting a single CPU usage number, and instead making throttled time, database connection pool utilization, and p99 latency all visible at once.",[16,233,234,235,238,239,244],{},"Building an observability stack from scratch on EKS or AKS costs time and money. With ",[34,236,74],{"href":72,"rel":237},[38],", a K3s-based service with Prometheus and Grafana included starting at ¥48,000\u002Fmonth, you can start using this kind of visibility right away. If you'd like to try running standard, vendor-lock-in-free Kubernetes without losing time to misdiagnosis, ",[34,240,243],{"href":241,"rel":242},"https:\u002F\u002Fwww.hexabase.com\u002Fcontact-us\u002F",[38],"get in touch with us"," to learn more.",{"title":25,"searchDepth":246,"depth":246,"links":247},2,[248,249,250,251,258],{"id":13,"depth":246,"text":14},{"id":78,"depth":246,"text":79},{"id":125,"depth":246,"text":126},{"id":160,"depth":246,"text":161,"children":252},[253,255,256,257],{"id":171,"depth":254,"text":172},3,{"id":183,"depth":254,"text":184},{"id":190,"depth":254,"text":191},{"id":203,"depth":254,"text":204},{"id":224,"depth":246,"text":225},"2026-07-28","When adding more Pods to Kubernetes doesn't fix latency, the real cause may not be CPU at all, but invisible throttling or database connection pool exhaustion. Here's how to diagnose and fix the hidden ceiling.","md","https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion\u002Feyecatch.webp","en",{},true,"\u002Fblog\u002Fen\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion",{"title":5,"description":260},"blog\u002Fen\u002Fkubernetes-cpu-throttling-connection-pool-exhaustion",[270,271,272,273,274,275],"kubernetes","k3s","performance-tuning","cpu-throttling","connection-pooling","hpa","kekSH6C924CWWCcHOcPXKWLYD-ycGBx2RuABc1xxd04",[278,287,296,305,313,322],{"path":279,"title":280,"description":281,"date":282,"tags":283},"\u002Fblog\u002Fen\u002Fkubernetes-gpu-multitenancy-namespace-vs-dedicated-node","Stop Letting One Team Hog Your Expensive GPUs: Why There's No Single Right Answer for Kubernetes Accelerator Sharing","Kubernetes GPU multi-tenancy isn't a binary choice between namespace isolation and dedicated nodes. This article breaks down the cost-vs-isolation trade-off and how to design a hybrid approach.","2026-08-13",[271,270,284,285,286],"gpu-multitenancy","cost-optimization","namespace-isolation",{"path":288,"title":289,"description":290,"date":291,"tags":292},"\u002Fblog\u002Fen\u002Fkubernetes-gitops-branch-antipattern-fleet-scaling","Your dev\u002Fstaging\u002Fprod Branches Are a Time Bomb: Why Kubernetes GitOps Really Breaks","Splitting dev\u002Fstaging\u002Fproduction by Git branch is a GitOps anti-pattern that undermines Kubernetes' declarative foundations. Learn why drift happens, how to migrate to a directory-based, trunk-based setup, and how to design for fleet-scale growth.","2026-08-10",[271,270,293,294,295],"gitops","argocd","devops",{"path":297,"title":298,"description":299,"date":300,"tags":301},"\u002Fblog\u002Fen\u002Fkubernetes-certificate-management-cert-manager-process-debt","The Cert Renewal Took One Line of Code and Two Months of Meetings: Why Kubernetes Certificate Management Is a Process Problem, Not a Technical One","Kubernetes certificate management is technically a matter of days. What actually takes time is the organizational process of getting sign-off. Here's how cert-manager automates the technical side, and how to design away the operational debt that remains.","2026-08-09",[271,270,302,303,304],"cert-manager","tls","security",{"path":306,"title":307,"description":308,"date":309,"tags":310},"\u002Fblog\u002Fen\u002Fai-agent-sandbox-kata-containers-kubernetes","AI Agent Code Isn't a \"Trusted Product\" Anymore. Kubernetes Sandbox Design Has an Answer","Code generated and executed by AI agents can no longer be treated as a trusted, reviewed product. This article explains the limits of container isolation and why Kata Containers' microVM isolation is becoming essential when designing AI agent sandboxes on Kubernetes.","2026-08-08",[271,270,311,312,304],"kata-containers","ai-agent",{"path":314,"title":315,"description":316,"date":317,"tags":318},"\u002Fblog\u002Fen\u002Fkubevirt-calico-live-migration-networking","Moving a VM Doesn't Have to Break the Connection: Inside KubeVirt and Calico's Live Migration Magic","Why doesn't live migrating a VM (KubeVirt) between Kubernetes nodes break the connection? We break down Calico's IP persistence and BGP route convergence, and what it means for teams moving off VMware.","2026-08-07",[271,270,319,320,321],"kubevirt","networking","managed-kubernetes",{"path":323,"title":324,"description":325,"date":326,"tags":327},"\u002Fblog\u002Fen\u002Fkubernetes-image-signing-sigstore-supply-chain","Anyone Can Rewrite an Image Tag. Why Kubernetes Needs Sigstore-Backed Signing to Prove Provenance","Container image signing explained: tags can be overwritten by anyone, and passing CI tests doesn't guarantee the image running in production is the one you built. Learn how Sigstore and Kyverno work together to reject unsigned images on Kubernetes\u002FK3s, integrated into a GitOps workflow.","2026-08-06",[271,270,328,293,304],"ci-cd",1786701442718]