Skip to content

Self-managed: no routing path for pod-addressed worker callbacks, so invocation-service and grpc-proxy are limited to one replica #2276

Description

@sparve-nv

Summary

Two invocation-plane services generate pod-addressed callback endpoints so a worker's response returns to the exact pod holding the request state. Self-managed deployments have no supported way to route those addresses from a worker, so both services must run a single replica.

What the services do today

Both build a callback authority by prefixing the dashed pod IP onto a base host.

invocation-service — src/invocation-plane-services/http-invocation/crates/server/src/worker_streams/mod.rs:

let pod_ip = pod_ip.replace(".", "-");
parts.authority = Some(Authority::try_from(format!("{}.{}", pod_ip, authority))?);

producing e.g. 10-0-1-5.invocation.nvcf.svc.cluster.local:8080. This is required for correctness: the request and response tokens live in pod-local DashMaps, so a callback landing on another replica cannot complete the waiting request.

grpc-proxy, HTTP/3 CONNECT path — src/invocation-plane-services/grpc-proxy/proxy/proxy.go:

parsedUrl.Host = strings.ReplaceAll(ip.To4().String(), ".", "-") + "." + parsedUrl.Host

producing e.g. 10-0-1-5.proxy.grpc.<domain>. Same purpose: the worker session is held by one proxy pod.

Both services also accept an explicit override — self_address for invocation-service, SELF_WORKER_FQDN for grpc-proxy — which takes priority and disables pod addressing.

The problem

When the compute plane is in a different network from the control plane, the worker cannot reach pod IPs or in-cluster pod DNS. The self-managed charts provide no routing path for the pod-addressed form either — no wildcard DNS, no pod-aware forwarding at the Gateway.

The only workable configuration is to set the explicit override to a public, load-balanced address. That is correct only while the service runs one replica: with more than one, the callback can be balanced to a pod that does not hold the state. The result is a dropped response and a client timeout, with no error at deploy time — the replica count is accepted and invocations simply start failing.

Service Pod-addressed form Self-managed configuration Result
invocation-service <dashed-pod-ip>.<service>:8080 public callback URL 1 replica
grpc-proxy <dashed-pod-ip>.<fqdn> public SELF_WORKER_FQDN, HTTP/1 1 replica

For grpc-proxy specifically, the pod-addressed form is only generated on the HTTP/3 CONNECT path (ENABLE_HTTP3_CONNECT). The HTTP/1 path used by self-managed deployments has no pod-addressed mode at all beyond a direct pod-IP fallback, which an external worker cannot reach. Note also that getConnectPaths offers HTTP/1 first and the worker takes the first config that connects, so HTTP/3 is a fallback rather than the primary transport.

This is already acknowledged in-tree

deploy/stacks/self-managed/global.yaml.gotmpl pins both services to a single replica whenever HA is enabled, with the same comment on each — :986 for invocation-service, :1105 for grpc-proxy:

{{- /* Single replica until Envoy; placement below is a no-op at 1 replica. */}}
replicaCount: 1

So the constraint is already understood and Envoy is already named as the resolution. This issue is asking for that work to be scheduled and its contract documented, rather than proposing a new design.

What we would like

  1. A supported routing path for pod-addressed worker callbacks in split control-plane / compute-plane deployments, so both services can run more than one replica. This is the core ask; the rest follows from it.
  2. Confirmation that the planned Envoy work covers both services. They carry the same comment but are different paths — invocation-service's response callback and grpc-proxy's CONNECT session. Worth stating explicitly whether one change addresses both, and roughly when it lands.
  3. State the supported replica count for invocation-service and grpc-proxy — the value supported today, and the target once Envoy lands.
  4. Document the TLS requirement, if pod-addressed hosts remain part of the answer. The pod-addressed host needs wildcard DNS and a matching wildcard SAN. setupTLS falls back to a locally generated certificate when PUBLIC_CERT_PATH / PRIVATE_KEY_PATH are unset, which will not validate against the worker's trust pool (src/libraries/go/worker/ca/ca.go), so mounting a real certificate would be required rather than optional. If the Envoy direction removes pod addressing, this becomes moot — either way it is worth confirming.

Why this matters

Without this, a split deployment has no invocation redundancy: a single invocation-service pod and a single grpc-proxy pod are both unprotected single points of failure on the request path, and neither can be rolled without an invocation outage.

Happy to contribute documentation or chart values.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions