본문으로 건너뛰기
← Systems Notebook / Ops
MLOps Notes02 / 15

Kubernetes 트래픽은 어디로 들어오는가 — kube-vip, MetalLB, Gateway API

온프레미스 RKE2에서 control-plane VIP와 워크로드 외부 IP를 구분하고, MetalLB와 NGINX Gateway Fabric으로 Gateway API 기반 진입 경로를 구성합니다.

이 글의 배경

HPE 엔지니어의 작업 자료를 바탕으로 정리한 글입니다. 개인 실험 기록이나 HPE 공식 문서와는 구분합니다. 본문의 환경과 버전은 글 작성 당시를 기준으로 합니다.

온프레미스 Kubernetes를 처음 구축할 때는 “클러스터는 정상인데 브라우저에서 서비스가 열리지 않는” 상황을 만날 수 있습니다. 클러스터 내부에서 Service와 Pod가 통신하더라도, 사내 네트워크의 사용자가 애플리케이션에 접속하려면 안정적인 외부 IP와 진입 경로가 필요하기 때문입니다.

control plane HA를 구성할 때도 주소의 역할을 구분해야 합니다. Kubernetes API에 접속할 주소, type: LoadBalancer 서비스에 할당할 주소, HTTP 요청을 서비스별로 나누는 규칙은 각각 해결하는 문제가 다릅니다.

이 글에서는 각 역할을 다음 구성요소에 맡기고, 단계별로 통신 경로를 확인합니다.

  • kube-vip: RKE2 server들 앞에서 Kubernetes API와 노드 등록 주소를 안정화합니다.
  • MetalLB: 클라우드 로드밸런서가 없는 온프레미스에서 LoadBalancer Service에 외부 IP를 할당하고 광고합니다.
  • NGINX Gateway Fabric: Gateway API를 구현해 hostname과 path에 따라 요청을 Kubernetes Service로 전달합니다.

마지막에는 이전 글에서 설치한 Rancher를 rancher.lab.example.com으로 노출합니다.

이 글에서 다루는 것

  • 노드 IP, control-plane VIP, Service 외부 IP가 왜 달라야 하는지
  • kube-vip의 ARP 리더 선출과 RKE2 static Pod 배치
  • MetalLB의 IPAddressPool과 L2Advertisement
  • GatewayClass, Gateway, HTTPRoute의 역할 분리
  • NGINX Gateway Fabric을 통한 데모 앱과 Rancher 노출
  • API VIP 페일오버, LoadBalancer IP, Route 상태의 검증 방법
  • ARP, 인증서, Route attachment 장애를 찾는 순서
  • IP 풀, namespace 허용 범위, TLS를 운영 환경에서 보호하는 방법

실습 환경과 치환값

아래 IP는 문서용 예약 대역입니다. 실제 실습에 사용할 주소는 네트워크 관리자에게 할당받아야 합니다. 같은 L2 네트워크에서 사용할 수 있고 DHCP 범위와 중복되지 않도록 예약한 주소여야 합니다.

의미 문서 예시 명령에서 사용할 이름
control-plane 노드 192.0.2.11~`192.0.2.13` NODE_1_IP~`NODE_3_IP`
control-plane VIP 192.0.2.100 CONTROL_PLANE_VIP
API DNS k8s-api.lab.example.com K8S_API_FQDN
MetalLB 풀 192.0.2.200~`192.0.2.219` LOADBALANCER_POOL
데모 애플리케이션 app.lab.example.com APP_FQDN
Rancher rancher.lab.example.com RANCHER_FQDN
kube-vip 이미지 v1.0.3 KUBE_VIP_VERSION
NGINX Gateway Fabric 2.6.6 NGF_VERSION

CONTROL_PLANE_VIP는 노드 NIC에 고정 설정하는 주소가 아닙니다. kube-vip 리더가 필요할 때만 소유합니다. MetalLB 풀도 노드 주소나 control-plane VIP와 겹치면 안 됩니다.

먼저 트래픽 경로를 분리해서 보자

단계별 흐름 / 01관리 요청과 애플리케이션 트래픽
관리 클라이언트kubectl·RKE2 agent
kube-vipCONTROL_PLANE_VIP
RKE2 server 1현재 VIP 소유 노드
01 / 04
관리 요청의 진입점

kubectl과 RKE2 agent는 관리용 VIP를 사용합니다. API는 6443, 노드 등록은 9345 포트입니다.

1

관리 클라이언트 → kube-vipAPI 6443 / 등록 9345

2

kube-vip → RKE2 server 1현재 VIP 경로

단계를 선택하면 자동 재생이 멈춥니다. 선의 번호와 아래 설명을 함께 읽어 주세요.

전체 단계 한눈에 읽기
  1. 관리 요청의 진입점

    kubectl과 RKE2 agent는 관리용 VIP를 사용합니다. API는 6443, 노드 등록은 9345 포트입니다.

    • 관리 클라이언트 → kube-vip: API 6443 / 등록 9345
    • kube-vip → RKE2 server 1: 현재 VIP 경로
  2. VIP 소유 노드가 바뀌면

    현재 리더가 멈추면 다른 노드가 VIP를 이어받습니다. server 2와 3은 순차 경유지가 아니라 대안입니다.

    • kube-vip → RKE2 server 1: 현재 VIP 경로
    • kube-vip → RKE2 server 2: 장애 시 VIP 이동
    • kube-vip → RKE2 server 3: 장애 시 VIP 이동
  3. 앱 요청은 다른 IP로

    외부 사용자는 MetalLB가 할당한 앱의 외부 IP로 접속합니다. 관리용 VIP와는 별개입니다.

    • 외부 사용자 → 앱의 외부 IP: HTTP/HTTPS
    • 앱의 외부 IP → Gateway 데이터 플레인: 요청 전달
  4. HTTPRoute로 앱 선택

    Gateway 데이터 플레인이 Gateway와 HTTPRoute 설정에 따라 Service를 선택하고 Pod로 요청을 전달합니다.

    • Gateway 데이터 플레인 → Kubernetes Service: Gateway + HTTPRoute
    • Kubernetes Service → Application Pods: 앱 요청
도식 원문
flowchart TB
  ADMIN["kubectl·RKE2 agent"] -->|"API 6443 / 등록 9345"| CPVIP["kube-vip\nCONTROL_PLANE_VIP"]
  CPVIP --> CP1["RKE2 server 1"]
  CPVIP -. "장애 시 VIP 이동" .-> CP2["RKE2 server 2"]
  CPVIP -. "장애 시 VIP 이동" .-> CP3["RKE2 server 3"]

  USER["외부 사용자"] -->|"HTTP/HTTPS"| LBIP["MetalLB가 할당한 외부 IP"]
  LBIP --> NGF["NGINX Gateway Fabric 데이터 플레인"]
  NGF -->|"Gateway + HTTPRoute"| SVC["Kubernetes Service"]
  SVC --> POD["Application Pods"]

control-plane VIP

RKE2 HA에서는 모든 server가 Kubernetes API를 제공할 수 있습니다. 하지만 클라이언트가 cp-01의 IP로만 접속하면 해당 노드에 장애가 났을 때 API에 접근하기 어렵습니다. RKE2 공식 HA 문서에서 “fixed registration address”를 요구하는 이유도 클라이언트와 추가 노드가 사용할 접속 주소를 안정적으로 유지하기 위해서입니다.

  • 6443/TCP: kubectl과 Kubernetes 구성요소가 사용하는 API
  • 9345/TCP: 새 RKE2 server/agent가 클러스터에 등록할 때 사용하는 supervisor

kube-vip는 같은 L2 네트워크의 control-plane 노드들 중 리더를 선출하고 VIP를 리더 NIC에 올립니다. 리더가 사라지면 다른 노드가 Gratuitous ARP를 보내 VIP의 MAC 주소가 바뀌었음을 네트워크에 알립니다.

Service 외부 IP

type: LoadBalancer는 “외부 IP를 원한다”는 Kubernetes API 선언입니다. 베어메탈에는 그 요청을 구현할 클라우드 프로바이더가 없으므로 아무 구성도 하지 않으면 EXTERNAL-IP가 <pending>에 머뭅니다.

MetalLB는 미리 승인된 풀에서 IP를 하나 할당합니다. L2 모드에서는 특정 노드가 해당 IP의 ARP 응답을 맡아 외부 트래픽을 받습니다. 이렇게 IP 할당과 네트워크 광고를 담당합니다.

L7 라우팅

외부 IP 하나를 얻었다고 app.lab.example.com과 rancher.lab.example.com 요청이 자동으로 나뉘지는 않습니다. NGINX Gateway Fabric은 Gateway API 리소스를 읽고 NGINX 데이터 플레인을 구성해 Host, path, header 같은 L7 조건으로 Service를 선택합니다.

왜 신규 구성은 Gateway API로 시작하는가

Ingress API는 Kubernetes에 남아 있지만 feature-frozen 상태입니다. 더 중요한 변화는 널리 사용되던 ingress-nginx 프로젝트가 2026년 3월에 유지보수를 중단하고 은퇴해 이후 릴리스, 버그 수정, 신규 보안 패치를 제공하지 않는다는 점입니다. 이 글에서는 ingress-nginx를 설치하지 않습니다.

Gateway API는 책임을 세 리소스로 나눕니다.

리소스 대표 소유자 결정하는 것
GatewayClass 인프라 제공자 어떤 controller 구현을 사용할지
Gateway 클러스터 운영자 listener, port, TLS, Route 허용 범위
HTTPRoute 애플리케이션 팀 hostname/path와 backend Service 매핑

역할을 나누면 애플리케이션 팀은 Gateway 자체를 수정할 권한 없이도 자기 namespace의 Route를 배포할 수 있습니다. 배포 후에는 Route의 status.conditions를 읽어 controller가 규칙을 수용했는지 확인할 수 있습니다.

실습 1: RKE2 control plane에 kube-vip 적용

이 절차는 이전 글의 첫 RKE2 server가 정상 실행 중인 상태에서 시작합니다. 3대 HA를 구성하려면 모든 control-plane 노드가 같은 L2 세그먼트에 있고, 동일한 이름의 NIC로 VIP를 광고할 수 있어야 합니다.

kube-vip를 처음 부팅부터 제공해야 하는 완전 자동화 환경에서는 static Pod 매니페스트와 kubeconfig 생성 순서의 경합을 별도 부트스트랩 코드로 해결해야 합니다. 이 실습은 첫 server를 먼저 시작한 뒤 VIP를 만들고, 추가 server를 VIP로 조인하는 이해하기 쉬운 순서를 사용합니다.

1. VIP와 인터페이스 확인

cp-01에서 실행합니다.

bash
export CONTROL_PLANE_VIP='192.0.2.100'
export K8S_API_FQDN='k8s-api.lab.example.com'
export VIP_INTERFACE='eth0'
export KUBE_VIP_VERSION='v1.0.3'

ip -br address show "${VIP_INTERFACE}"
ping -c 2 "${CONTROL_PLANE_VIP}" || true

VIP를 활성화하기 전인데 ping 응답이 온다면 이미 다른 장비가 그 주소를 사용 중일 수 있습니다. 중복 여부를 확인하기 전에는 계속 진행하지 않습니다. VIP_INTERFACE는 실제 노드의 인터페이스 이름으로 바꿉니다.

DNS에서는 k8s-api.lab.example.com이 control-plane VIP를 가리키도록 준비합니다. 첫 글의 /etc/rancher/rke2/config.yaml에도 이 FQDN이 tls-san으로 들어 있어야 합니다.

bash
sudo grep -A3 '^tls-san:' /etc/rancher/rke2/config.yaml

2. kube-vip 매니페스트 생성

RKE2가 제공하는 containerd에 kube-vip 이미지를 pull하고, 이미지의 manifest generator를 실행합니다. --services와 --enableLoadBalancer는 넣지 않습니다. Service 외부 IP는 MetalLB가 전담하고, kube-vip 자체의 IPVS load-balancing 기능도 이 구성에서는 사용하지 않습니다.

bash
export CTR='/var/lib/rancher/rke2/bin/ctr'
export CONTAINERD_ADDRESS='/run/k3s/containerd/containerd.sock'
export KUBE_VIP_IMAGE="ghcr.io/kube-vip/kube-vip:${KUBE_VIP_VERSION}"

sudo "${CTR}" \
  --address "${CONTAINERD_ADDRESS}" \
  --namespace k8s.io \
  image pull "${KUBE_VIP_IMAGE}"

sudo install -d -m 0755 \
  /var/lib/rancher/rke2/agent/pod-manifests

sudo "${CTR}" \
  --address "${CONTAINERD_ADDRESS}" \
  --namespace k8s.io \
  run --rm --net-host \
  "${KUBE_VIP_IMAGE}" kube-vip-manifest \
  /kube-vip manifest pod \
  --interface "${VIP_INTERFACE}" \
  --address "${CONTROL_PLANE_VIP}" \
  --vipSubnet 32 \
  --controlplane \
  --arp \
  --leaderElection \
  --namespace kube-system | \
  sudo tee \
    /var/lib/rancher/rke2/agent/pod-manifests/kube-vip.yaml \
    >/dev/null

공식 generator의 static Pod 예시는 kubeadm의 /etc/kubernetes/admin.conf를 마운트합니다. RKE2의 관리자 kubeconfig 경로는 /etc/rancher/rke2/rke2.yaml이므로 host path와 container mount path를 RKE2 경로로 바꿉니다.

bash
sudo sed -i \
  's#/etc/kubernetes/admin.conf#/etc/rancher/rke2/rke2.yaml#g' \
  /var/lib/rancher/rke2/agent/pod-manifests/kube-vip.yaml

sudo grep -nE \
  'image:|vip_interface|address|cp_enable|svc_enable|rke2.yaml' \
  /var/lib/rancher/rke2/agent/pod-manifests/kube-vip.yaml

검토 결과에서 다음을 확인합니다.

  • 이미지 태그가 의도한 KUBE_VIP_VERSION입니다.
  • cp_enable은 true입니다.
  • svc_enable은 false이거나 존재하지 않습니다.
  • 인터페이스와 주소가 정확합니다.
  • kubeconfig hostPath가 RKE2 경로인지 확인합니다.
  • NET_ADMIN, NET_RAW capability와 hostNetwork: true가 있습니다.

RKE2 kubelet이 static Pod 파일을 감지하면 kube-vip가 시작됩니다.

bash
kubectl -n kube-system get pod -o wide | grep kube-vip
ip address show dev "${VIP_INTERFACE}" | grep "${CONTROL_PLANE_VIP}"
curl --cacert /var/lib/rancher/rke2/server/tls/server-ca.crt \
  "https://${K8S_API_FQDN}:6443/livez" || true

마지막 요청에서는 API 인증이 필요해 HTTP 인증 오류가 반환될 수 있습니다. 이 요청으로 확인하려는 범위는 DNS·TCP 연결과 TLS 인증서 검증까지입니다. API가 실제로 요청을 처리할 준비가 되었는지는 kubectl get --raw='/readyz'로 확인합니다.

3. 추가 server를 VIP로 조인

cp-02, cp-03의 보호된 설정 파일은 다음 의미를 가집니다.

yaml
server: https://k8s-api.lab.example.com:9345
token: ${RKE2_TOKEN}
tls-san:
  - k8s-api.lab.example.com
ingress-controller: none

여기서 ${RKE2_TOKEN}은 예시 표기입니다. 첫 server의 /var/lib/rancher/rke2/server/node-token을 비밀 전달 경로로 주입해야 하며 Git에 저장하지 않습니다. 추가 server가 Ready가 된 후 같은 kube-vip 매니페스트를 각 노드의 RKE2 static Pod 경로에 배치합니다.

bash
kubectl get nodes -o wide
kubectl -n kube-system get pods -o wide | grep kube-vip
kubectl -n kube-system get lease | grep kube-vip

3대가 모두 Ready이고 각 control-plane 노드에 kube-vip static Pod mirror가 보이며, 리더 Lease가 하나 존재해야 합니다.

4. 페일오버 검증

운영 중인 노드를 갑자기 내리기 전에 현재 VIP 소유자를 기록합니다.

bash
for node in cp-01 cp-02 cp-03; do
  echo "=== ${node} ==="
  ssh "${node}" \
    "ip address show dev ${VIP_INTERFACE} | grep ${CONTROL_PLANE_VIP} || true"
done

승인된 점검 시간에 VIP 리더의 RKE2 서비스를 중지하고, 별도 관리 터미널에서 API 요청이 다시 성공하는 시간을 측정합니다.

bash
while true; do
  date '+%H:%M:%S'
  kubectl --request-timeout=2s get --raw='/readyz' || true
  sleep 1
done

새 리더가 VIP를 인계하면 API 요청도 회복되어야 합니다. ping 응답만으로는 Kubernetes API HA가 동작한다고 판단할 수 없으므로 요청의 회복까지 확인합니다. 테스트가 끝나면 중지한 RKE2 서비스를 다시 시작하고 모든 노드와 etcd 상태를 확인합니다.

실습 2: MetalLB로 LoadBalancer IP 제공

1. 설치

bash
helm repo add metallb https://metallb.github.io/metallb
helm repo update metallb

helm upgrade --install metallb metallb/metallb \
  --namespace metallb-system \
  --create-namespace \
  --version 0.16.1 \
  --wait --timeout 10m

MetalLB speaker는 네트워크 광고를 위해 높은 권한이 필요합니다. Pod Security Admission을 강하게 적용한 클러스터라면 공식 설치 문서의 namespace 정책을 검토합니다.

bash
kubectl -n metallb-system get deployment,daemonset,pod
kubectl -n metallb-system get crd | grep metallb || true

controller와 각 노드의 speaker가 준비된 뒤 IP 풀을 만듭니다. CRD가 준비되기 전에 바로 적용하면 webhook 연결 오류가 날 수 있습니다.

2. IPAddressPool과 L2Advertisement

bash
cat > metallb-l2.yaml <<'EOF'
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
  name: platform-pool
  namespace: metallb-system
spec:
  addresses:
    - 192.0.2.200-192.0.2.219
  autoAssign: true
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
  name: platform-l2
  namespace: metallb-system
spec:
  ipAddressPools:
    - platform-pool
EOF

kubectl apply -f metallb-l2.yaml
kubectl -n metallb-system get ipaddresspool,l2advertisement

인터페이스를 명시하지 않으면 MetalLB가 노드의 적절한 인터페이스에서 광고합니다. 멀티 NIC 호스트에서는 예상하지 않은 망으로 ARP가 나가지 않도록 L2Advertisement.spec.interfaces를 명시하고 실제 패킷 캡처로 확인합니다.

3. 단순 LoadBalancer Service로 먼저 검증

Gateway를 구성하기 전에 가장 작은 Service로 MetalLB 경로를 시험합니다. 여기서 외부 IP 할당과 접속을 확인해 두면, 이후 Gateway를 구성할 때 문제가 생겨도 점검할 범위를 좁힐 수 있습니다.

bash
kubectl create namespace lb-test
kubectl -n lb-test create deployment echo \
  --image=registry.k8s.io/e2e-test-images/agnhost:2.53 \
  -- /agnhost netexec --http-port=8080
kubectl -n lb-test expose deployment echo \
  --name echo \
  --port 80 \
  --target-port 8080 \
  --type LoadBalancer
kubectl -n lb-test get service echo -w

정상이라면 EXTERNAL-IP에 MetalLB 풀의 주소 하나가 나타납니다. 같은 L2 네트워크의 클라이언트에서 요청합니다.

bash
export ECHO_LB_IP="$(kubectl -n lb-test get service echo \
  -o jsonpath='{.status.loadBalancer.ingress[0].ip}')"
curl "http://${ECHO_LB_IP}/hostname"

Pod hostname이 응답되면 IP 할당, ARP 광고, Service와 endpoint 경로가 모두 동작합니다. 검증 후 테스트 자원을 지웁니다.

bash
kubectl delete namespace lb-test

실습 3: NGINX Gateway Fabric 설치

1. Gateway API CRD 설치

NGINX Gateway Fabric이 지원하는 CRD 버전을 맞춰 설치합니다. 아래 v2.6.6 ref는 NGINX 공식 설치 문서가 해당 NGF 버전에 사용하는 구성을 가리킵니다.

bash
export NGF_VERSION='2.6.6'

kubectl kustomize \
  "https://github.com/nginx/nginx-gateway-fabric/config/crd/gateway-api/standard?ref=v${NGF_VERSION}" | \
  kubectl apply -f -

kubectl get crd gatewayclasses.gateway.networking.k8s.io
kubectl get crd gateways.gateway.networking.k8s.io
kubectl get crd httproutes.gateway.networking.k8s.io

이 글은 Standard channel만 사용합니다. 실험 기능이 필요한 경우 구현체 호환성과 승격·제거 정책을 별도로 검토합니다.

2. NGF controller 설치

bash
helm upgrade --install ngf \
  oci://ghcr.io/nginx/charts/nginx-gateway-fabric \
  --namespace nginx-gateway \
  --create-namespace \
  --version "${NGF_VERSION}" \
  --set nginxGateway.productTelemetry.enable=false \
  --wait --timeout 10m
bash
kubectl -n nginx-gateway get deployment,pod
kubectl get gatewayclass nginx

다음 단계로 넘어가기 전에 controller가 준비되었는지, nginx GatewayClass의 Accepted 상태가 True인지 확인합니다. 이후 Gateway를 생성하면 NGF가 Gateway별 NGINX 데이터 플레인과 기본 LoadBalancer Service를 만듭니다. MetalLB는 이 Service에 외부 IP를 할당합니다.

실습 4: Gateway와 HTTPRoute로 앱 노출

1. 테스트용 TLS 인증서

아래 self-signed 인증서는 실습에서 경로를 확인하기 위한 것입니다. 운영 환경에서는 사내 CA 또는 신뢰할 수 있는 ACME 발급자와 cert-manager를 사용하고, 개인키 생성·보관·갱신 책임을 명확히 합니다.

bash
openssl req -x509 -nodes -newkey rsa:3072 -days 30 \
  -keyout /tmp/lab-wildcard.key \
  -out /tmp/lab-wildcard.crt \
  -subj '/CN=*.lab.example.com' \
  -addext 'subjectAltName=DNS:*.lab.example.com'

kubectl -n nginx-gateway create secret tls lab-wildcard-tls \
  --cert=/tmp/lab-wildcard.crt \
  --key=/tmp/lab-wildcard.key \
  --dry-run=client -o yaml | kubectl apply -f -

rm -f /tmp/lab-wildcard.key /tmp/lab-wildcard.crt

개인키 임시 파일은 즉시 지웁니다. 이 파일을 쉘 기록, Git, 공유 폴더에 남기지 않습니다.

2. Route를 허용할 namespace 표시

Gateway에 모든 namespace의 Route를 무제한 허용하지 않고 label selector를 사용합니다.

bash
kubectl create namespace demo --dry-run=client -o yaml | kubectl apply -f -
kubectl label namespace demo gateway-access=true --overwrite
kubectl label namespace cattle-system gateway-access=true --overwrite

3. Gateway 생성

bash
cat > platform-gateway.yaml <<'EOF'
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: platform-gateway
  namespace: nginx-gateway
spec:
  gatewayClassName: nginx
  listeners:
    - name: http
      protocol: HTTP
      port: 80
      hostname: "*.lab.example.com"
      allowedRoutes:
        namespaces:
          from: Selector
          selector:
            matchLabels:
              gateway-access: "true"
    - name: https
      protocol: HTTPS
      port: 443
      hostname: "*.lab.example.com"
      tls:
        mode: Terminate
        certificateRefs:
          - kind: Secret
            name: lab-wildcard-tls
      allowedRoutes:
        namespaces:
          from: Selector
          selector:
            matchLabels:
              gateway-access: "true"
EOF

kubectl apply -f platform-gateway.yaml
kubectl -n nginx-gateway get gateway platform-gateway

PROGRAMMED=True가 되고 ADDRESS에 MetalLB 풀의 IP가 나타나야 합니다.

bash
kubectl -n nginx-gateway describe gateway platform-gateway
kubectl -n nginx-gateway get service \
  -l gateway.networking.k8s.io/gateway-name=platform-gateway

4. 데모 Deployment, Service, HTTPRoute

bash
cat > demo-app.yaml <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata:
  name: demo-app
  namespace: demo
spec:
  replicas: 2
  selector:
    matchLabels:
      app: demo-app
  template:
    metadata:
      labels:
        app: demo-app
    spec:
      containers:
        - name: app
          image: registry.k8s.io/e2e-test-images/agnhost:2.53
          args: ["netexec", "--http-port=8080"]
          ports:
            - name: http
              containerPort: 8080
          readinessProbe:
            httpGet:
              path: /healthz
              port: http
---
apiVersion: v1
kind: Service
metadata:
  name: demo-app
  namespace: demo
spec:
  selector:
    app: demo-app
  ports:
    - name: http
      port: 80
      targetPort: http
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: demo-app
  namespace: demo
spec:
  parentRefs:
    - name: platform-gateway
      namespace: nginx-gateway
      sectionName: http
  hostnames:
    - app.lab.example.com
  rules:
    - matches:
        - path:
            type: PathPrefix
            value: /
      backendRefs:
        - name: demo-app
          port: 80
EOF

kubectl apply -f demo-app.yaml
kubectl -n demo rollout status deployment/demo-app --timeout=5m
kubectl -n demo get httproute demo-app

DNS 변경 전에도 curl --resolve로 정확한 Host 헤더와 IP를 시험할 수 있습니다.

bash
export GATEWAY_IP="$(kubectl -n nginx-gateway get gateway platform-gateway \
  -o jsonpath='{.status.addresses[0].value}')"

curl --resolve "app.lab.example.com:80:${GATEWAY_IP}" \
  http://app.lab.example.com/hostname

Pod hostname이 반환되면 외부 IP → NGF 데이터 플레인 → HTTPRoute → Service → Pod로 이어지는 요청 경로가 동작하는지 확인할 수 있습니다.

실습 5: Rancher를 HTTPS Route로 연결

이전 글에서 Rancher Chart를 ingress.enabled=false로 설치했습니다. Rancher Service는 cattle-system에 있으므로 같은 namespace에 HTTPRoute를 만듭니다.

bash
cat > rancher-route.yaml <<'EOF'
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: rancher
  namespace: cattle-system
spec:
  parentRefs:
    - name: platform-gateway
      namespace: nginx-gateway
      sectionName: https
  hostnames:
    - rancher.lab.example.com
  rules:
    - matches:
        - path:
            type: PathPrefix
            value: /
      backendRefs:
        - name: rancher
          port: 80
EOF

kubectl apply -f rancher-route.yaml
kubectl -n cattle-system get httproute rancher
kubectl -n cattle-system describe httproute rancher

self-signed 인증서를 사용했기 때문에 아래 실습 요청만 --insecure로 보냅니다. 운영 환경의 검증 절차에는 이 옵션을 넣지 않습니다.

bash
curl --insecure --head \
  --resolve "rancher.lab.example.com:443:${GATEWAY_IP}" \
  https://rancher.lab.example.com/

HTTP 응답이 오면 DNS에 rancher.lab.example.com을 GATEWAY_IP로 연결하고 브라우저에서 확인합니다. TLS 경고 없이 운영하려면 신뢰 체인이 있는 인증서로 교체해야 합니다.

검증 체크리스트와 예상 결과

bash
# control plane 고정 주소
getent hosts k8s-api.lab.example.com
kubectl get --raw='/readyz'
kubectl -n kube-system get pods -o wide | grep kube-vip

# MetalLB
kubectl -n metallb-system get pods
kubectl -n metallb-system get ipaddresspool,l2advertisement

# Gateway API와 NGF
kubectl get gatewayclass nginx
kubectl -n nginx-gateway get gateway platform-gateway
kubectl -n demo get httproute demo-app
kubectl -n cattle-system get httproute rancher

완료 기준은 다음과 같습니다.

  • API DNS의 인증서와 tls-san이 맞고 /readyz가 성공합니다.
  • kube-vip static Pod가 각 control-plane 노드에 있고 리더가 VIP를 소유합니다.
  • Gateway의 address가 MetalLB 풀 안에 있으며 Programmed=True입니다.
  • 두 HTTPRoute의 Accepted, ResolvedRefs 조건이 True입니다.
  • curl --resolve로 데모 앱과 Rancher 모두 응답합니다.
  • VIP 리더 장애 시 제한된 중단 후 API 요청이 다시 성공합니다.

장애 대응

kube-vip Pod가 시작되지 않는다

bash
kubectl -n kube-system describe pod -l component=kube-vip
sudo grep -n 'rke2.yaml' \
  /var/lib/rancher/rke2/agent/pod-manifests/kube-vip.yaml
sudo journalctl -u rke2-server -n 200 --no-pager

가장 먼저 이미지 pull, kubeconfig hostPath, NIC 이름, capability를 봅니다. static Pod는 일반 Deployment처럼 API에서 삭제해도 원본 파일이 남아 있으면 kubelet이 다시 만듭니다. 실제 선언은 각 노드의 /var/lib/rancher/rke2/agent/pod-manifests/kube-vip.yaml입니다.

VIP는 보이는데 API 인증서 오류가 난다

k8s-api.lab.example.com 또는 VIP가 RKE2의 tls-san에 없을 가능성이 큽니다.

bash
openssl s_client \
  -connect k8s-api.lab.example.com:6443 \
  -servername k8s-api.lab.example.com \
  </dev/null 2>/dev/null | \
  openssl x509 -noout -subject -issuer -ext subjectAltName

SAN을 고친 뒤 공식 인증서 갱신 절차를 따릅니다. TLS 검증을 끄는 것은 해결이 아닙니다.

LoadBalancer Service가 <pending>이다

bash
kubectl -n metallb-system get pods
kubectl -n metallb-system get ipaddresspool,l2advertisement -o yaml
kubectl describe service -A | grep -A8 -B3 LoadBalancer
kubectl -n metallb-system logs deployment/metallb-controller --tail=200

MetalLB가 IP를 할당하려면 풀과 광고 객체가 모두 있어야 합니다. 풀 고갈, 잘못된 namespace, webhook 미준비, 다른 controller와의 충돌을 확인합니다.

외부 IP는 있지만 같은 LAN에서 접속되지 않는다

L2Advertisement는 라우터를 새로 만들지 않습니다. 클라이언트와 광고 노드가 L2로 닿는지, 스위치의 ARP 보안·포트 보안이 Gratuitous ARP를 막는지 확인합니다.

bash
ip neigh show | grep "${GATEWAY_IP}"
sudo tcpdump -ni "${VIP_INTERFACE}" "arp or host ${GATEWAY_IP}"

멀티랙·라우팅 경계 환경에서는 L2보다 BGP가 맞을 수 있습니다. BGP는 라우터 ASN, peer, 필터, ECMP까지 네트워크 팀과 공동 설계합니다.

HTTPRoute가 Accepted=False다

bash
kubectl -n demo describe httproute demo-app
kubectl -n nginx-gateway describe gateway platform-gateway
kubectl -n demo get service,endpointslices
  • parentRefs.namespace와 sectionName이 정확한지 확인합니다.
  • Route namespace에 gateway-access=true label이 있는지 확인합니다.
  • hostname이 Gateway listener의 wildcard 범위에 들어오는지 확인합니다.
  • backend Service 이름과 port가 같은 namespace에서 해석되는지 확인합니다.
  • 다른 namespace의 Secret이나 Service를 참조한다면 ReferenceGrant가 필요한지 확인합니다.

Gateway는 Programmed인데 address가 없다

NGF가 만든 데이터 플레인 Service의 타입과 MetalLB 이벤트를 함께 봅니다.

bash
kubectl -n nginx-gateway get service -o wide
kubectl -n nginx-gateway get pod -o wide
kubectl -n nginx-gateway logs deployment/ngf-nginx-gateway-fabric --tail=200

controller 자체의 내부 Service와 Gateway 데이터 플레인의 외부 Service를 혼동하지 않습니다.

운영과 보안 고려사항

kube-vip와 MetalLB의 주소 영역을 분리한다

이 실습에서는 kube-vip가 control plane을, MetalLB가 Service를 담당하도록 기능을 나눕니다. 같은 IP나 같은 Service를 두 controller가 관리하면 리더십과 ARP 광고가 충돌할 수 있기 때문입니다. 주소 할당표에도 owner, 용도, DNS, 장애 전환 방식을 함께 기록해 관리 주체를 구분합니다.

IP 풀은 네트워크 자산이다

MetalLB 풀은 DHCP 범위에서 제외하고 IPAM에 등록합니다. 자동 할당이 위험한 운영 서비스는 별도의 IPAddressPool과 autoAssign: false 정책을 사용해 명시적으로 선택합니다. 개발·운영 풀을 분리하면 실수로 운영 IP를 소비하는 일을 줄일 수 있습니다.

allowedRoutes: All을 기본값으로 두지 않는다

어떤 namespace가 공용 Gateway에 Route를 연결할 수 있는지 label과 RBAC로 제한합니다. 같은 hostname을 여러 팀이 선점하는 경우도 admission policy나 Git 리뷰에서 차단합니다. Gateway 운영자와 앱 팀의 권한을 분리하는 것이 Gateway API의 장점을 살리는 방법입니다.

TLS 개인키의 소유권을 정한다

이 글에서 사용하는 self-signed wildcard는 실습용입니다. 운영 환경에서는 wildcard의 사용 범위와 인증서 발급자를 정하고, Secret 접근 권한과 갱신 실패 경보도 준비해야 합니다. Gateway namespace에 접근할 수 있는 주체는 해당 도메인의 트래픽을 가로챌 수 있으므로, 권한을 정할 때 이 범위까지 고려합니다.

kube-proxy IPVS와 kube-vip를 혼동하지 않는다

Kubernetes 1.35의 deprecation 대상은 kube-proxy의 ipvs 프록시 모드입니다. kube-vip의 ARP 기반 주소 소유권과는 다른 개념입니다. 이 실습에서는 kube-vip의 --enableLoadBalancer도 사용하지 않으며, RKE2 kube-proxy는 기본값을 유지합니다.

RKE2에서 nftables를 선택하려면 최소 v1.36.3+rke2r1, v1.35.7+rke2r1, v1.34.10+rke2r1 계열의 2026년 7월 릴리스부터라는 버전 게이트와 experimental 상태를 확인하고 CNI별 설정을 함께 검증합니다.

요약

  • control-plane VIP, LoadBalancer Service IP, L7 Route는 각각 해결하는 문제가 다릅니다.
  • kube-vip는 RKE2 API와 9345 등록 주소를 하나의 안정된 IP로 만듭니다.
  • MetalLB는 베어메탈에서 LoadBalancer Service의 IP를 할당하고 L2 또는 BGP로 광고합니다.
  • NGINX Gateway Fabric은 Gateway API를 구현해 hostname/path 요청을 Service로 전달합니다.
  • 이 실습의 신규 구성에서는 은퇴한 ingress-nginx 설치 경로 대신 Gateway API를 사용합니다.
  • 상태 확인은 ping 하나가 아니라 API readiness, Service address, Gateway Programmed, HTTPRoute Accepted/ResolvedRefs, 실제 HTTP 요청까지 이어져야 합니다.

이전 글: 온프레미스 Kubernetes를 왜 RKE2로 구축하는가
다음 글: Kubernetes 스토리지 입문 — PV/PVC부터 NFS CSI와 RWX까지

공식 참고자료