Skip to content
Documentation Background

Service Discovery Deep Dive

  • Service discovery is how one microservice finds and communicates with another - by name, not by IP. Kubernetes automates the entire name-to-IP translation chain so application code only ever needs to know the target Service name.
  • Division of labour: developers configure which Service names to call; Kubernetes handles ClusterIP allocation, DNS registration, EndpointSlice tracking, and kernel-level routing on every node.
  • Foundation: every Service gets a stable ClusterIP and a DNS entry. The DNS entry is what makes discovery reliable regardless of Pod churn underneath.

Kubernetes Service Discovery

When app-a needs to talk to app-b, the flow is:

  1. Name configured (manual): the developer sets the target Service name (app-b) in app-a’s config - this is the only manual step.
  2. DNS query (automated): app-a’s container queries the cluster DNS for app-b’s IP.
  3. ClusterIP returned (automated): CoreDNS resolves the name to app-b’s virtual ClusterIP.
  4. Traffic sent (automated): app-a sends packets to the ClusterIP; kube-proxy kernel rules intercept and forward to a healthy Pod.

The Routing Challenge: ClusterIPs are Non-Routable

Section titled “The Routing Challenge: ClusterIPs are Non-Routable”

ClusterIPs live on a virtual service network that has no routes on containers or nodes. Packets fall through to default gateways:

flowchart LR
    subgraph SourcePod["📦 SOURCE CONTAINER (Pod Netns)"]
        direction TB
        P1["1️⃣ Packet sent<br/><code>dst: ClusterIP</code>"]
        P2["2️⃣ Route lookup<br/><i>No route in netns</i>"]
        P3["3️⃣ Default gateway<br/><code>via veth interface</code>"]
        P1 --> P2 --> P3
    end

    subgraph HostNode["🖥️ HOST NODE (Kernel & kube-proxy)"]
        direction TB
        H1["4️⃣ Host interface<br/><i>Packet arrives</i>"]
        H2["5️⃣ Routing check<br/><i>Towards default gw</i>"]
        H3["6️⃣ Kernel netfilter<br/><i>Processes packet</i>"]
        H4["7️⃣ kube-proxy rules<br/><code>iptables / IPVS</code>"]
        H5["8️⃣ DNAT rewrite<br/><code>ClusterIP ➔ Pod IP</code>"]
        H1 --> H2 --> H3 --> H4 --> H5
    end

    subgraph DestPod["🎯 TARGET BACKEND POD"]
        direction TB
        D1["9️⃣ Delivered to Pod<br/><code>dst: Pod IP:targetPort</code>"]
    end

    SourcePod -->|"leaves container netns"| HostNode
    HostNode -->|"routed via CNI"| DestPod

kube-proxy watches the API server for Service and EndpointSlice changes and pre-programs these kernel rules on every node. Traffic interception is transparent to the application.

PhaseResponsible entityAction
Configure target nameDeveloperHardcode or inject Service name into app config
DNS requestApplication containerQuery nameserver from /etc/resolv.conf
Name resolutionCoreDNSResolve FQDN to ClusterIP
Packet interceptionNode kernel + kube-proxyRewrite destination IP to Pod IP
Kubernetes Service Discovery

On Pod creation, kubelet populates /etc/resolv.conf in every container:

search dev.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.96.0.10
options ndots:5
  • nameserver: the ClusterIP of the kube-dns Service - all DNS queries go here.
  • search: ordered list of suffixes appended to unqualified (short) names:
    • <current-namespace>.svc.cluster.local - tried first
    • svc.cluster.local
    • cluster.local
  • options ndots:5: if a name has fewer than 5 dots, search suffixes are tried before treating it as an absolute FQDN.
Name typeExampleScopeHow it resolves
Short namewebSame namespace onlyResolver appends <local-ns>.svc.cluster.local automatically
FQDNweb.prod.svc.cluster.localAny namespaceSent directly without search suffix appending

Short name resolution (intra-namespace): when a Pod in dev runs curl web:8080:

  1. Resolver appends dev.svc.cluster.local - queries web.dev.svc.cluster.local
  2. CoreDNS returns the ClusterIP for web in dev
  3. kube-proxy rules forward traffic to a healthy Pod

Cross-namespace resolution: to reach web in prod from dev, the FQDN is required:

Terminal window
curl web.prod.svc.cluster.local:8080

Short names always expand to the local namespace first - they cannot reach remote namespaces.

Terminal window
kubectl exec -it <pod-name> -n <namespace> -- cat /etc/resolv.conf
# Expected: search <ns>.svc.cluster.local svc.cluster.local cluster.local
# nameserver 10.96.0.10
# options ndots:5

The default ndots:5 causes up to 5 sequential DNS lookups for any external FQDN that has fewer than 5 dots. For example, calling api.github.com from a Pod in the default namespace triggers:

1. api.github.com.default.svc.cluster.local <- NXDOMAIN
2. api.github.com.svc.cluster.local <- NXDOMAIN
3. api.github.com.cluster.local <- NXDOMAIN
4. api.github.com. <- RESOLVED (absolute)

This adds latency on every external call and hammers CoreDNS with wasted queries. Two mitigations:

  • Trailing dot trick: use api.github.com. (note the trailing dot) in your config. The resolver treats a trailing dot as an absolute FQDN and skips search suffixes entirely.
  • Lower ndots via dnsConfig: override per-Pod in the spec (covered in Services - DNS Policy):
spec:
dnsConfig:
options:
- name: ndots
value: "2" # external FQDNs with 2+ dots resolve directly

Every Service gets a Fully Qualified Domain Name:

<service-name>.<namespace>.svc.cluster.local

Examples:

  • web in default - web.default.svc.cluster.local
  • web in dev - web.dev.svc.cluster.local
  • web in prod - web.prod.svc.cluster.local
  • Service names must be unique within a namespace - one web per namespace.
  • Service names can be identical across namespaces - web in dev and web in prod are separate, independent Services.
  • This enables parallel environments (dev/staging/prod) on the same cluster with identical manifests and no naming collisions.

Regular Services resolve to the ClusterIP (load-balanced across all Pods). Headless Services for StatefulSets expose each Pod individually:

<pod-name>.<headless-svc-name>.<namespace>.svc.cluster.local

For a StatefulSet db backed by headless Service db-headless in namespace data:

Terminal window
# Individual Pod addresses - stable across restarts:
db-0.db-headless.data.svc.cluster.local # pod 0
db-1.db-headless.data.svc.cluster.local # pod 1
db-2.db-headless.data.svc.cluster.local # pod 2
# The headless service itself returns all Pod IPs as A records (no ClusterIP):
nslookup db-headless.data.svc.cluster.local
# returns: 10.244.1.5, 10.244.2.3, 10.244.3.7

This is how StatefulSet Pods find each other by stable identity - for clustered databases (PostgreSQL replicas, Kafka brokers, etcd peers) where a Pod must know the address of a specific peer. See DaemonSets and StatefulSets for the full StatefulSet coverage.

Practical Example: Same Name, Different Namespaces

Section titled “Practical Example: Same Name, Different Namespaces”

Two deployments - web Service in dev (responds “Hello from DEV”) and web Service in prod (responds “Hello from PROD”):

Terminal window
# From a Pod in the dev namespace:
curl web:8080
# Hello from DEV <- resolves to dev namespace via search suffix
curl web.prod.svc.cluster.local:8080
# Hello from PROD <- FQDN bypasses local search, targets prod directly
Terminal window
# Confirm the two services have different ClusterIPs:
kubectl get svc web -n dev # e.g. 10.96.138.186
kubectl get svc web -n prod # e.g. 10.96.147.32

Every cluster runs an internal cluster DNS (CoreDNS) as the service registry - a Kubernetes-native application built from standard API objects, all in kube-system with label k8s-app=kube-dns.

Kubernetes CoreDNS Workflow
ComponentObject nameRoleKey details
Deploymentcoredns (or kube-dns)Manages DNS Pod lifecycle and replicasTypically 2+ replicas for redundancy
Podscoredns-xxxxx-xxxxxProcess A/AAAA and SRV resolution queriesImage: registry.k8s.io/coredns/coredns
Servicekube-dnsStable ClusterIP frontend for all DNS queriesAlways named kube-dns across all clusters; ports 53/UDP, 53/TCP, 9153/TCP
EndpointSlicekube-dns-xxxxxTracks active CoreDNS Pod IPsPrefixed kube-dns- (e.g. kube-dns-jb72g)
Terminal window
# All 4 components share this label
kubectl get deploy,pods,svc,endpointslice -n kube-system -l k8s-app=kube-dns
# Or individually:
kubectl get deploy -n kube-system -l k8s-app=kube-dns
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl get svc kube-dns -n kube-system
kubectl get endpointslice -n kube-system -l k8s-app=kube-dns

CoreDNS is configured via a Corefile stored in the coredns ConfigMap in kube-system. You can edit it to add forwarding rules, stub zones, hosts overrides, or custom caching.

Terminal window
# View the current Corefile
kubectl get configmap coredns -n kube-system -o yaml

Default Corefile structure:

.:53 {
errors
health {
lameduck 5s
}
ready
kubernetes cluster.local in-addr.arpa ip6.arpa { # handles cluster DNS
pods insecure
fallthrough in-addr.arpa ip6.arpa
ttl 30
}
prometheus :9153 # metrics endpoint
forward . /etc/resolv.conf { # upstream for external names
max_concurrent 1000
}
cache 30 # cache TTL in seconds
loop
reload
loadbalance
}

Forward a specific domain to an internal resolver (split-horizon DNS for on-prem):

corp.example.com:53 {
errors
forward . 192.168.1.53 # on-prem DNS resolver
cache 30
}

Add static host entries (useful for legacy services without DNS):

hosts {
10.0.1.50 legacy-db.internal
fallthrough
}

Override external DNS (redirect a domain to a cluster Service for testing):

rewrite name api.external.com internal-api-svc.default.svc.cluster.local

After editing, CoreDNS picks up changes automatically via the reload plugin - no restart needed. Validate with:

Terminal window
# Edit the ConfigMap
kubectl edit configmap coredns -n kube-system
# Watch CoreDNS logs for reload confirmation
kubectl logs -f <coredns-pod> -n kube-system
# [INFO] plugin/reload: Running configuration SHA512 = <new-hash>

When a developer posts a Service manifest, 8 automated stages fire in sequence:

flowchart TD
    subgraph Row1["Phase 1: Registration & Storage (Control Plane)"]
        direction LR
        S1["1️⃣ POST YAML<br/><code>API server</code>"] --> S2["2️⃣ ClusterIP allocated<br/><i>Service IP range</i>"] --> S3["3️⃣ Persisted to etcd<br/><i>Cluster store</i>"] --> S4["4️⃣ CoreDNS detects<br/><i>via API watch</i>"]
    end

    subgraph Row2["Phase 2: DNS & Data Plane Sync (Worker Nodes)"]
        direction RL
        S8["8️⃣ IPVS rules live<br/><i>Kernel interception</i>"]
        S7["7️⃣ kube-proxy syncs<br/><i>Node daemon</i>"] --> S8
        S6["6️⃣ EndpointSlices created<br/><i>Controller</i>"] --> S7
        S5["5️⃣ DNS records written<br/><i>A/AAAA & SRV</i>"] --> S6
    end

    Row1 -->|"API watch events trigger sync"| Row2
  • Stage 1 - POST to API server: manifest authenticated, RBAC checked, schema validated.
  • Stage 2 - ClusterIP allocation: a unique virtual IP is drawn from the service network range.
  • Stage 3 - Persisted to etcd: Service state and ClusterIP written to the cluster store.
  • Stage 4 - CoreDNS detects: CoreDNS maintains a watch stream on the API server; new Service detected instantly.
  • Stage 5 - DNS records written: CoreDNS generates A/AAAA records (name to ClusterIP) and SRV records (named ports to port numbers).
  • Stage 6 - EndpointSlices created: the EndpointSlice controller matches the Service selector against healthy Pods and populates slices with Pod IPs.
  • Stage 7 - kube-proxy syncs: the kube-proxy daemon on every node pulls the new Service and EndpointSlice data.
  • Stage 8 - Kernel rules programmed: kube-proxy writes IPVS/iptables rules to intercept ClusterIP traffic and load-balance it to Pod IPs.
Kubernetes Service Discovery

As clusters scale to fleet management, cross-cluster discovery is standardized by the Multi-Cluster Services (MCS) API. Rather than using messy API federations or bespoke syncing controllers, MCS uses ServiceExport and ServiceImport to define cross-cluster boundaries declaratively.

  • ClusterSet: A fleet of clusters managed together (the boundary of trust).
  • ServiceExport: An opt-in resource created in the source cluster to declare a local Service should be accessible to the rest of the ClusterSet.
  • ServiceImport: Created automatically in destination clusters by an MCS controller (e.g., Cilium Cluster Mesh or GKE MCS) to mirror the exported Service.
# Source Cluster: Export the local 'payments' service
apiVersion: multicluster.x-k8s.io/v1alpha1
kind: ServiceExport
metadata:
name: payments
namespace: finance

Once exported, the service becomes globally discoverable across the ClusterSet using a specialized DNS domain:

Terminal window
# Destination Cluster: resolve the imported service
nslookup payments.finance.svc.clusterset.local

The <svc>.<ns>.svc.clusterset.local FQDN pattern is standard. When a Pod queries this domain, CoreDNS resolves it to a virtual IP (VIP) managed by the MCS controller, which securely tunnels the traffic (often via eBPF/WireGuard tunnels or service meshes) to the remote endpoints.


In large clusters, every DNS query from every Pod hits the CoreDNS Pods directly, creating a bottleneck. NodeLocal DNSCache solves this by running a DNS caching agent as a DaemonSet on every node - queries are answered locally without a network hop to CoreDNS.

  • How it works: a node-local-dns DaemonSet binds to a link-local IP (typically 169.254.20.10) on each node. kubelet configures new Pods to use this local IP as nameserver instead of the kube-dns ClusterIP.
  • Cache hits: served in microseconds from the node, no CoreDNS round-trip.
  • Cache misses: the node agent forwards to CoreDNS and caches the response for subsequent queries.
  • Conntrack benefit: uses TCP for upstream queries to CoreDNS, avoiding conntrack race conditions that can drop UDP DNS responses under load.
Terminal window
# Check if NodeLocal DNSCache is deployed
kubectl get daemonset -n kube-system node-local-dns
# Verify the link-local IP is listening on a node
kubectl debug node/<node-name> -it --image=busybox -- nslookup kubernetes 169.254.20.10

ExternalDNS is a controller that watches Kubernetes Services and Ingresses and automatically creates/updates records in external DNS providers (Route53, Cloud DNS, Azure DNS, Cloudflare, etc.). It bridges the gap between in-cluster DNS and public DNS.

  1. Annotate a Service or Ingress with the desired hostname:
apiVersion: v1
kind: Service
metadata:
name: web-public
annotations:
external-dns.alpha.kubernetes.io/hostname: app.example.com
spec:
type: LoadBalancer
  1. ExternalDNS detects the annotation, reads the Service’s external IP (assigned by the cloud load balancer), and creates an A record in your DNS provider:
app.example.com -> 203.0.113.45 # created automatically in Route53 / Cloud DNS
  1. When the Service is deleted or the annotation removed, ExternalDNS deletes the record.
  • No manual DNS management - hostnames follow the lifecycle of Kubernetes objects.
  • Works with Ingress - can create DNS records for every Ingress host rule automatically.
  • Provider-agnostic - same annotations work across AWS, GCP, Azure, and 30+ other providers.
Terminal window
# Check ExternalDNS is running and watching the correct sources
kubectl get pods -n kube-system -l app.kubernetes.io/name=external-dns
kubectl logs -n kube-system -l app.kubernetes.io/name=external-dns | grep "Desired change"
# Desired change: CREATE app.example.com A [203.0.113.45] in zone Z1234

Terminal window
# Deployment should show all replicas ready
kubectl get deploy -n kube-system -l k8s-app=kube-dns
# NAME READY UP-TO-DATE AVAILABLE
# coredns 2/2 2 2
# Pods should be Running with 0 restarts
kubectl get pods -n kube-system -l k8s-app=kube-dns
# Check logs for config errors or plugin failures
kubectl logs <coredns-pod-name> -n kube-system
# Healthy output:
# .:53
# [INFO] plugin/reload: Running configuration SHA512 = 591cf3...
# CoreDNS-1.11.1 linux/arm64
Terminal window
kubectl get svc kube-dns -n kube-system
# NAME TYPE CLUSTER-IP PORT(S)
# kube-dns ClusterIP 10.96.0.10 53/UDP,53/TCP,9153/TCP

Cross-check that 10.96.0.10 matches the nameserver in a container’s /etc/resolv.conf:

Terminal window
kubectl exec -it <any-pod> -- cat /etc/resolv.conf

Step 3 - Verify EndpointSlice is Populated

Section titled “Step 3 - Verify EndpointSlice is Populated”
Terminal window
kubectl get endpointslice -n kube-system -l k8s-app=kube-dns
# NAME ADDRESSTYPE PORTS ENDPOINTS
# kube-dns-jb72g IPv4 9153,53,53 10.244.1.9,10.244.1.14

Empty ENDPOINTS means CoreDNS Pods are not passing readiness - check Pod status and logs.

Launch a diagnostic Pod with DNS utilities:

Terminal window
kubectl run dnsutils --image=registry.k8s.io/e2e-test-images/jessie-dnsutils:1.7 -it --rm

Inside the Pod, query the kubernetes Service (present in every cluster):

Terminal window
nslookup kubernetes
# Server: 10.96.0.10
# Address: 10.96.0.10#53
#
# Name: kubernetes.default.svc.cluster.local
# Address: 10.96.0.1
  • Lines 1-2: DNS queries are reaching kube-dns at port 53 - CoreDNS is reachable.
  • Lines 3-4: name resolved to FQDN and correct ClusterIP - DNS records are intact.

Cross-verify the returned IP:

Terminal window
kubectl get svc kubernetes # ClusterIP should match the Address above

If nslookup: can't resolve kubernetes - delete the Pods; the Deployment recreates them:

Terminal window
kubectl delete pod -n kube-system -l k8s-app=kube-dns
# Wait, then verify and re-test
kubectl get pods -n kube-system -l k8s-app=kube-dns
SymptomLikely causeFix
nslookup: can't resolve <name>CoreDNS crash-looping or nameserver IP mismatchRestart CoreDNS Pods; verify ClusterIP matches resolv.conf
DNS resolves but traffic failskube-proxy rules stale or EndpointSlice emptyCheck kubectl get endpointslice; restart kube-proxy DaemonSet
Short name resolves wrong ServicePod in wrong namespaceUse FQDN to explicitly target the correct namespace
Cross-namespace comms failUsing short name for cross-namespace callReplace with <svc>.<ns>.svc.cluster.local
Terminal window
# Delete the diagnostic Pod (--rm handles this if used above, otherwise:)
kubectl delete pod dnsutils