How to Deploy a Next.js Application on Kubernetes: A Practical Guide
I moved my first Next.js app to Kubernetes because Vercel's bill for a client project quietly crossed four figures a month — mostly from image optimization and edge function invocations we didn't fully control. Kubernetes wasn't cheaper on day one. It got cheaper once we understood our own traffic patterns and stopped paying for someone else's abstraction. That's really the honest reason teams end up here: not because Kubernetes is trendy, but because at some point you want the knobs yourself.
This guide walks through actually shipping a Next.js app on a production Kubernetes cluster — building the right kind of minimal Docker image, writing manifests that won't bite you at 2 a.m., handling environment variables properly, and scaling without guessing. No hand-waving, no "it depends" without explaining the trade-offs.
Why Bother with Kubernetes for Next.js?
Platforms like Vercel and Netlify exist precisely so you don't have to manage raw compute. For a marketing site or a small SaaS with predictable traffic, they're the right call — genuinely, I still reach for managed platforms first for anything under a certain scale. Kubernetes earns its operational complexity when one or more of these conditions are met:
- Co-located Microservices: Your infrastructure already lives in Kubernetes (internal APIs, databases, message queues), and you want your Next.js frontend in the same virtual private cloud (VPC) sharing service discovery and secrets.
- Strict Data Compliance: Governance rules mandate that compute and network traffic stay within specific geographic boundaries or dedicated bare-metal clusters.
- Fine-Grained Autoscaling: You need custom scaling policies based on in-house metrics (e.g., custom Prometheus request latency) rather than platform black-box triggers.
- Cost Optimization at Scale: Owning dedicated nodes becomes substantially cheaper than per-invocation serverless pricing when traffic is consistently high and sustained.
If none of those apply, stay on a managed platform. But if your team needs complete infrastructure control, let's configure this properly.
Step 1: Containerize Next.js the Right Way (Standalone Mode)
The single biggest mistake is copying a generic Node.js Dockerfile and wondering why the image weighs 1.2GB. Next.js includes a built-in standalone output feature designed specifically for containerized workloads.
First, enable it in next.config.js (or next.config.mjs):
/** @type {import('next').NextConfig} */
const nextConfig = {
output: 'standalone',
reactStrictMode: true,
};
module.exports = nextConfig;
This instructs Next.js to trace package dependencies and bundle only the strictly required production files into .next/standalone. In our production builds, this reduced image footprints from 1.1GB down to roughly 180MB.
Here is the hardened multi-stage Dockerfile:
# --- Stage 1: Install Dependencies ---
FROM node:20-alpine AS deps
RUN apk add --no-cache libc6-compat
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
# --- Stage 2: Build Application ---
FROM node:20-alpine AS builder
WORKDIR /app
COPY --from=deps /app/node_modules ./node_modules
COPY . .
ENV NEXT_TELEMETRY_DISABLED=1
RUN npm run build
# --- Stage 3: Production Runtime ---
FROM node:20-alpine AS runner
WORKDIR /app
ENV NODE_ENV=production
ENV NEXT_TELEMETRY_DISABLED=1
# Run as non-root user for Pod Security Standards
RUN addgroup --system --gid 1001 nodejs && \
adduser --system --uid 1001 nextjs
COPY --from=builder /app/public ./public
COPY --from=builder --chown=nextjs:nodejs /app/.next/standalone ./
COPY --from=builder --chown=nextjs:nodejs /app/.next/static ./.next/static
USER nextjs
EXPOSE 3000
ENV PORT=3000
ENV HOSTNAME="0.0.0.0"
CMD ["node", "server.js"]
next/image extensively, add npm install sharp in the runner stage. The standalone Node.js runtime defaults to the slower JS image processor if the native libvips binding isn't present.Build and push using semantic tags instead of latest:
docker build -t yourregistry/nextjs-app:1.0.0 .
docker push yourregistry/nextjs-app:1.0.0
Step 2: Production Kubernetes Manifests
To run reliably, your workload requires a Deployment with proper health checks, a Service for internal routing, and an Ingress controller for TLS termination.
1. The Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: nextjs-app
labels:
app: nextjs-app
spec:
replicas: 3
selector:
matchLabels:
app: nextjs-app
template:
metadata:
labels:
app: nextjs-app
spec:
containers:
- name: nextjs-app
image: yourregistry/nextjs-app:1.0.0
imagePullPolicy: IfNotPresent
ports:
- containerPort: 3000
env:
- name: NODE_ENV
value: "production"
- name: API_URL
valueFrom:
configMapKeyRef:
name: nextjs-config
key: api_url
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
cpu: "500m"
memory: "512Mi"
readinessProbe:
httpGet:
path: /api/health
port: 3000
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 2
livenessProbe:
httpGet:
path: /api/health
port: 3000
initialDelaySeconds: 15
periodSeconds: 20
timeoutSeconds: 3
/): Do not point your liveness/readiness probes to the root route. If home page rendering waits on a slow database query or an external CRM, Kubernetes will falsely flag your pod as dead and trigger cascading restarts across the cluster.Create a dedicated health endpoint in Next.js App Router (app/api/health/route.ts):
export async function GET() {
return Response.json({ status: 'ok', uptime: process.uptime() });
}
2. The Service
apiVersion: v1
kind: Service
metadata:
name: nextjs-service
spec:
type: ClusterIP
selector:
app: nextjs-app
ports:
- protocol: TCP
port: 80
targetPort: 3000
3. The Ingress
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: nextjs-ingress
annotations:
cert-manager.io/cluster-issuer: "letsencrypt-prod"
nginx.ingress.kubernetes.io/proxy-body-size: "10m"
spec:
ingressClassName: nginx
tls:
- hosts:
- example.com
secretName: nextjs-tls
rules:
- host: example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: nextjs-service
port:
number: 80
Managing Environment Variables Without Leaking Secrets
Next.js behaves differently with environment variables depending on their prefix:
NEXT_PUBLIC_*: Inlined directly into the client-side JavaScript bundle duringnpm run build. Changing these in a Kubernetes ConfigMap post-build will have zero effect.- Standard Server Variables: Read at runtime by Node.js. These map seamlessly to Kubernetes Secrets and ConfigMaps.
Build-time Strategy
Inject NEXT_PUBLIC_ variables as Docker build arguments (ARG) inside your CI/CD pipeline, building unique images per staging environment.
Runtime Config API
Keep public configs in a /api/config endpoint fetched at browser initialization. Allows promoting the exact same immutable container image across Dev, QA, and Prod.
For sensitive backend keys, use standard Secrets referenced via secretKeyRef:
apiVersion: v1
kind: Secret
metadata:
name: nextjs-secrets
type: Opaque
stringData:
DATABASE_URL: "postgresql://db_user:password_xyz@postgres-svc:5432/production_db"
Autoscaling via Horizontal Pod Autoscaler (HPA)
Because Server-Side Rendering (SSR) in Node.js is single-threaded per event loop, CPU spikes are the primary indicator of saturation. Configure an HPA target based on average CPU utilization:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: nextjs-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: nextjs-app
minReplicas: 3
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65
Architecture Comparison: Where Should You Run?
| Evaluation Metric | Kubernetes (Self-Hosted / Managed EKS/GKE) | Vercel / Netlify | Standalone VM / Docker Swarm |
|---|---|---|---|
| Operational Overhead | High (Requires ongoing cluster & ingress management) | Virtually Zero | Medium |
| Cold Starts | None (Always-on containers) | Possible on rarely visited edge routes | None |
| Cost Predictability | Predictable compute billing based on node sizes | Variable based on function invocations & bandwidth | Low and fixed |
| Private Networking | Native cluster VPC peering & internal DNS | Requires enterprise secure gateways | Manual VPN configuration |
| Ideal Use Case | High-traffic enterprise apps with internal APIs | Fast-growth SaaS, documentation, marketing | Early-stage MVPs, internal dev tools |
Production Traps & How to Avoid Them
Frequently Asked Questions
Do I need a separate Dockerfile for development and production?
Not necessarily. Most teams use next dev locally without containers to retain instant Hot Module Replacement (HMR), keeping the multi-stage Dockerfile dedicated to CI/CD production pipelines.
Can I run Next.js API routes on Kubernetes, or do I need a separate backend?
Next.js API/Route Handlers run seamlessly inside the standalone Node.js container. You only need to extract them if specific endpoints require dedicated GPU resources, long-running background tasks, or decoupled scaling schedules.
How do I achieve zero-downtime rolling updates?
Kubernetes uses a RollingUpdate strategy by default. As long as your readinessProbe is accurate and you configure an adequate terminationGracePeriodSeconds (typically 30s) to allow ongoing connections to drain, updates will deploy without dropping traffic.
How does Next.js Image Optimization behave in a container?
The standalone container includes the image optimization handler. However, running sharp image transforms under high traffic consumes substantial CPU. Offload caching to a CDN like Cloudflare or CloudFront placed in front of your Ingress controller.
Comments