[{"data":1,"prerenderedAt":293},["ShallowReactive",2],{"blog-en-kubernetes-certificate-management-cert-manager-process-debt":3,"blog-related-en-kubernetes-certificate-management-cert-manager-process-debt":241,"blog-en-kubernetes-certificate-management-cert-manager-process-debt-alt":230},{"id":4,"title":5,"author":6,"body":7,"date":224,"description":225,"extension":226,"image":227,"locale":228,"meta":229,"navigation":230,"path":231,"seo":232,"stem":233,"tags":234,"__hash__":240},"blog\u002Fblog\u002Fen\u002Fkubernetes-certificate-management-cert-manager-process-debt.md","The Cert Renewal Took One Line of Code and Two Months of Meetings: Why Kubernetes Certificate Management Is a Process Problem, Not a Technical One","Kubo Team",{"type":8,"value":9,"toc":212},"minimark",[10,14,25,34,39,46,49,58,61,65,71,85,94,109,114,117,121,127,130,139,148,157,166,170,176,179,185,192,196,199],[11,12,13],"p",{},"Mention \"Kubernetes certificate management\" to most engineers and they picture a solved problem: just install cert-manager. But in practice, certificate renewals turn into fires for reasons that are rarely technical. Writing the Terraform or Helm code to automate a renewal might take one or two days — getting permission to ship it to production can take one or two months.",[11,15,16,17,24],{},"Certificate-related incidents are still one of the most common categories of outage. According to ",[18,19,23],"a",{"href":20,"rel":21},"https:\u002F\u002Fgartsolutions.com\u002Fpreventing-downtime-from-expired-certificates\u002F",[22],"nofollow","a report from Gart Solutions",", more than 70% of organizations have experienced at least one certificate-related outage in the past year. If you've ever chased down a mysterious production outage and found an expired TLS certificate at the bottom of it, you already know the pattern.",[11,26,27,28,33],{},"This article breaks down why certificate renewal so often becomes a process problem rather than a technical one, covers what cert-manager does and doesn't automate, and outlines how to design certificate operations so the debt never accumulates in the first place. Some of that debt can be avoided entirely by choosing a managed Kubernetes service like ",[18,29,32],{"href":30,"rel":31},"https:\u002F\u002Fkubo.hexabase.io\u002F",[22],"Kubo",", which sidesteps the problem structurally.",[35,36,38],"h2",{"id":37},"_1-the-real-reason-certificates-expire-isnt-technical-its-procedural","1. The Real Reason Certificates Expire Isn't Technical — It's Procedural",[11,40,41],{},[42,43],"img",{"alt":44,"src":45},"section01","https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-certificate-management-cert-manager-process-debt\u002Fsection01.webp",[11,47,48],{},"In many organizations, the certificate renewal flow is surprisingly analog. Someone notices a certificate is nearing expiration, manually creates a CSR (certificate signing request), sends it to another team or an external CA contact, waits for the signed certificate to come back, and finally uploads it by hand. Multiple approval gates and cross-team handoffs are baked into this process, and a single renewal can easily take weeks.",[11,50,51,52,57],{},"The deeper problem is that ownership of this procedure is often left ambiguous. When no one clearly owns a certificate, no one notices the warning signs, and the failure only surfaces once production is already down. ",[18,53,56],{"href":54,"rel":55},"https:\u002F\u002Fwww.encryptionconsulting.com\u002F10-cases-of-certificate-outages-involving-human-error\u002F",[22],"A roundup from Encryption Consulting"," describes how a 2018 certificate management failure at Ericsson knocked roughly 32 million people off 4G service in the UK alone, and how, in the 2017 Equifax breach, an expired certificate on a monitoring device went unnoticed for 19 months. In both cases, the technical root cause was simple — an expired certificate — but what turned it into a disaster was the absence of any mechanism to catch it.",[11,59,60],{},"What makes this worse is that the technical fix (say, writing Terraform code to automate renewal) is often finished in a short amount of time, while the explanation and consensus-building needed to convince stakeholders that \"nothing will break\" and \"this fits our existing workflow\" takes far longer than expected. Writing code has become one of the easier parts of the job. The hard part is designing the process for safely rolling out a change.",[35,62,64],{"id":63},"_2-what-cert-manager-automates-and-what-it-doesnt","2. What cert-manager Automates — and What It Doesn't",[11,66,67],{},[42,68],{"alt":69,"src":70},"section02","https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-certificate-management-cert-manager-process-debt\u002Fsection02.webp",[11,72,73,74,79,80,84],{},"The Kubernetes ecosystem has a standard tool for reducing this operational burden: cert-manager. According to the ",[18,75,78],{"href":76,"rel":77},"https:\u002F\u002Fcert-manager.io\u002Fdocs\u002Fusage\u002Fingress\u002F",[22],"official cert-manager documentation",", cert-manager ships with a subcomponent called \"ingress-shim\" that automatically generates a matching Certificate resource whenever you add specific annotations to an Ingress resource. Add a single annotation like ",[81,82,83],"code",{},"cert-manager.io\u002Fcluster-issuer",", and cert-manager takes over issuing, storing, and renewing the certificate.",[11,86,87,88,93],{},"Certificate issuers are managed through two resource types: Issuer and ClusterIssuer. Per the ",[18,89,92],{"href":90,"rel":91},"https:\u002F\u002Fcert-manager.io\u002Fdocs\u002Fconcepts\u002Fissuer\u002F",[22],"official cert-manager documentation on Issuers",", an Issuer is namespace-scoped and can only be referenced by resources in that same namespace, while a ClusterIssuer is cluster-scoped and can be referenced by Ingress and Certificate resources across multiple namespaces. If you want a single, unified issuance policy across the whole organization, use a ClusterIssuer; if different teams need different CAs, use Issuers.",[11,95,96,97,102,103,108],{},"Let's Encrypt is the most common certificate issuer in this setup, but its production ACME endpoint has rate limits. According to ",[18,98,101],{"href":99,"rel":100},"https:\u002F\u002Fletsencrypt.org\u002Fdocs\u002Frate-limits\u002F",[22],"Let's Encrypt's official rate limits documentation",", a single account is capped at 300 new orders every 3 hours, and any one domain is capped at 50 certificates every 7 days. Hitting these limits during development — by repeatedly requesting certificates against the production endpoint — is common enough that Let's Encrypt strongly recommends using its ",[18,104,107],{"href":105,"rel":106},"https:\u002F\u002Fletsencrypt.org\u002Fdocs\u002Fstaging-environment\u002F",[22],"staging environment"," instead. This is one reason cert-manager can't just be \"installed and forgotten\": you still need to understand these operational constraints well enough to design your Issuer setup and monitoring around them.",[110,111,113],"h3",{"id":112},"the-expiry-risk-cert-manager-alone-cant-prevent","The Expiry Risk cert-manager Alone Can't Prevent",[11,115,116],{},"cert-manager handles automatic renewal, but it doesn't guarantee anyone notices when a renewal fails. Failed ACME challenges, DNS misconfigurations, permission issues that block a Secret update — failures happen even inside an automated pipeline. This \"automated but unmonitored\" state is exactly the gap the next section addresses.",[35,118,120],{"id":119},"_3-turning-certificate-renewal-from-a-ritual-into-a-non-event","3. Turning Certificate Renewal from a Ritual Into a Non-Event",[11,122,123],{},[42,124],{"alt":125,"src":126},"section03","https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-certificate-management-cert-manager-process-debt\u002Fsection03.webp",[11,128,129],{},"Truly freeing certificate management from operational debt requires more than installing cert-manager — it requires designing the renewal process so it no longer depends on a human noticing anything. The most effective way to do that is to continuously monitor certificate expiry as a metric.",[11,131,132,133,138],{},"According to the ",[18,134,137],{"href":135,"rel":136},"https:\u002F\u002Fprometheus.io\u002Fdocs\u002Fintroduction\u002Foverview\u002F",[22],"official Prometheus documentation",", Prometheus is an open-source monitoring tool that collects and stores numeric time-series data, evaluates rules against those metrics, and fires notifications through Alertmanager when conditions are met. By feeding cert-manager's metrics into Prometheus and alerting when a certificate's remaining validity drops below a threshold, you shift the operating assumption from \"someone notices\" to \"the system tells you.\"",[11,140,141,142,147],{},"Certificate lifetimes themselves are also getting shorter, and that trend can't be ignored. According to a ",[18,143,146],{"href":144,"rel":145},"https:\u002F\u002Fwww.sectigo.com\u002Fblog\u002F200-day-ssl-certificate-expiration-risk",[22],"Sectigo blog post",", the industry-standard maximum certificate validity period is scheduled to drop from 398 days to 200 days starting in March 2026, then to 100 days in 2027, and finally to 47 days by 2029. Shorter validity periods mean more frequent renewals, and the burden of manual operations scales up dramatically as a result. The same article warns that at that point, \"manual certificate management becomes unsustainable.\" In other words, designing renewal automation and monitoring now is really about front-loading a cost that will only grow later.",[11,149,150,151,156],{},"Coverage of this topic is growing in Japanese-language media as well. ",[18,152,155],{"href":153,"rel":154},"https:\u002F\u002Fatmarkit.itmedia.co.jp\u002Fait\u002Farticles\u002F2410\u002F25\u002Fnews013.html",[22],"An ITmedia article"," walks through introducing cert-manager to prevent the classic \"forgot to renew\" incident, but like most introductory guides, it focuses on installation rather than on clarifying process ownership or building in monitoring and alerting — the areas that actually reduce operational debt.",[11,158,159,160,165],{},"The same design principles apply in Rancher-based environments. ",[18,161,164],{"href":162,"rel":163},"https:\u002F\u002Foneuptime.com\u002Fblog\u002Fpost\u002F2026-03-20-cert-manager-in-rancher\u002Fview",[22],"A technical blog post on running cert-manager under Rancher"," walks through deploying cert-manager via Helm on K3s\u002FRancher and integrating it with Let's Encrypt or a Rancher CA. Even on a lightweight K3s cluster, the importance of automating and monitoring certificate management doesn't change.",[35,167,169],{"id":168},"_4-with-managed-kubernetes-this-whole-process-never-has-to-exist","4. With Managed Kubernetes, This Whole Process Never Has to Exist",[11,171,172],{},[42,173],{"alt":174,"src":175},"section04","https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-certificate-management-cert-manager-process-debt\u002Fsection04.webp",[11,177,178],{},"As we've seen, turning certificate management into something you can trust to just work requires stacking up multiple layers: not just installing cert-manager, but designing your Issuer setup, testing against rate limits, and building out Prometheus-based monitoring and alerting. That's not a small amount of work, and many organizations get stuck at this stage and just resign themselves to \"certificate renewal is a pain.\"",[11,180,181,184],{},[18,182,32],{"href":30,"rel":183},[22]," is a K3s-based managed Kubernetes service that ships with cert-manager and Let's Encrypt built in from the start. That means the Issuer design and monitoring setup this article describes is already done by the time you start using it. Compared to building, validating, and monitoring cert-manager yourself on EKS or AKS, this significantly cuts both the upfront build cost and the risk of the setup depending on one specific person's knowledge.",[11,186,187,188,191],{},"If the procedural complexity of certificate renewal is a recurring headache, it's worth considering an option that removes the procedure entirely. ",[18,189,32],{"href":30,"rel":190},[22]," delivers full Kubernetes functionality without making you carry that operational debt from day one.",[35,193,195],{"id":194},"_5-summary","5. Summary",[11,197,198],{},"The core challenge in certificate management isn't whether you have cert-manager — it's whether anyone owns the renewal process and how it's monitored. Installing an automation tool and eliminating procedural debt are related but distinct problems. cert-manager itself is excellent, but without designing your Issuer\u002FClusterIssuer setup, operating within rate limits, and building in monitoring and alerting, you can still end up in the position of \"we installed it, and it expired anyway without anyone noticing.\"",[11,200,201,202,205,206,211],{},"As certificate lifecycles keep shrinking, the cost of deferring this operational debt keeps rising every year. If you're tired of a meeting breaking out every time a certificate needs renewing, start by clarifying who owns your certificate operations and how they're monitored. From there, choosing a managed Kubernetes platform like ",[18,203,32],{"href":30,"rel":204},[22]," that ships with certificate management built in is one solid option. If you'd like to talk through your certificate operations, feel free to reach out via our ",[18,207,210],{"href":208,"rel":209},"https:\u002F\u002Fwww.hexabase.com\u002Fcontact-us\u002F",[22],"contact page",".",{"title":213,"searchDepth":214,"depth":214,"links":215},"",2,[216,217,221,222,223],{"id":37,"depth":214,"text":38},{"id":63,"depth":214,"text":64,"children":218},[219],{"id":112,"depth":220,"text":113},3,{"id":119,"depth":214,"text":120},{"id":168,"depth":214,"text":169},{"id":194,"depth":214,"text":195},"2026-08-09","Kubernetes certificate management is technically a matter of days. What actually takes time is the organizational process of getting sign-off. Here's how cert-manager automates the technical side, and how to design away the operational debt that remains.","md","https:\u002F\u002Fcdn.kubo.hexabase.io\u002Fimages\u002Fblog\u002Fkubernetes-certificate-management-cert-manager-process-debt\u002Feyecatch.webp","en",{},true,"\u002Fblog\u002Fen\u002Fkubernetes-certificate-management-cert-manager-process-debt",{"title":5,"description":225},"blog\u002Fen\u002Fkubernetes-certificate-management-cert-manager-process-debt",[235,236,237,238,239],"k3s","kubernetes","cert-manager","tls","security","_131Z2aOaKmMlNZWnCxKle_ikmsp5nXn5RyRXC8-Wb4",[242,250,258,267,276,285],{"path":243,"title":244,"description":245,"date":246,"tags":247},"\u002Fblog\u002Fen\u002Fai-agent-sandbox-kata-containers-kubernetes","AI Agent Code Isn't a \"Trusted Product\" Anymore. Kubernetes Sandbox Design Has an Answer","Code generated and executed by AI agents can no longer be treated as a trusted, reviewed product. This article explains the limits of container isolation and why Kata Containers' microVM isolation is becoming essential when designing AI agent sandboxes on Kubernetes.","2026-08-08",[235,236,248,249,239],"kata-containers","ai-agent",{"path":251,"title":252,"description":253,"date":254,"tags":255},"\u002Fblog\u002Fen\u002Fkubernetes-image-signing-sigstore-supply-chain","Anyone Can Rewrite an Image Tag. Why Kubernetes Needs Sigstore-Backed Signing to Prove Provenance","Container image signing explained: tags can be overwritten by anyone, and passing CI tests doesn't guarantee the image running in production is the one you built. Learn how Sigstore and Kyverno work together to reject unsigned images on Kubernetes\u002FK3s, integrated into a GitOps workflow.","2026-08-06",[235,236,256,257,239],"ci-cd","gitops",{"path":259,"title":260,"description":261,"date":262,"tags":263},"\u002Fblog\u002Fen\u002Fkubernetes-secrets-rbac-etcd-encryption","Base64 Isn't Encryption: Why Kubernetes Secrets Pass Right Through, and the RBAC Design Traps That Make It Worse","Kubernetes Secrets are only Base64-encoded, not encrypted. Learn how plaintext-equivalent storage in etcd and over-permissioned RBAC lead to real incidents, plus the concrete Secrets management practices you need for production K3s.","2026-07-29",[236,235,264,265,239,266],"secrets-management","rbac","etcd-encryption",{"path":268,"title":269,"description":270,"date":271,"tags":272},"\u002Fblog\u002Fen\u002Fshadow-ai-kubernetes-admission-control-governance","Rogue Deployments Are Wrecking Your Company: The Shadow AI Problem Inside Kubernetes Clusters, and Admission Control as the Fix","Shadow AI isn't just unauthorized SaaS tools. It's happening inside your Kubernetes clusters too. Here's the risk it creates, and how Admission Control turns detection into real governance.","2026-07-17",[236,235,273,274,275,239],"shadow-ai","admission-control","kyverno",{"path":277,"title":278,"description":279,"date":280,"tags":281},"\u002Fblog\u002Fen\u002Fkubernetes-v136-k3s-managed-cost-reduction","Managed Kubernetes is Too Expensive. The Reality of 'Full K8s Operations Under $400\u002FMonth' with K3s Lightweight and v1.36 Security Enhancements","Explore how to leverage Kubernetes v1.36 'Haru' enhanced User Namespaces and security features in K3s lightweight environments. Discover managed K3s operational strategies and 2026 infrastructure selection guidelines that achieve 60% cost reduction compared to EKS.","2026-05-28",[235,236,282,283,284,239],"kubernetes-v136","managed-kubernetes","cost-optimization",{"path":286,"title":287,"description":288,"date":289,"tags":290},"\u002Fblog\u002Fen\u002Fkubernetes-gpu-multitenancy-namespace-vs-dedicated-node","Stop Letting One Team Hog Your Expensive GPUs: Why There's No Single Right Answer for Kubernetes Accelerator Sharing","Kubernetes GPU multi-tenancy isn't a binary choice between namespace isolation and dedicated nodes. This article breaks down the cost-vs-isolation trade-off and how to design a hybrid approach.","2026-08-13",[235,236,291,284,292],"gpu-multitenancy","namespace-isolation",1786701442583]