RunReconciler now runs two tickers:
- 30s: retry Failed namespaces (existing behavior)
- 60s: DispatchResync on Active namespaces; since registerNamespace is
idempotent this is a no-op when SA/RoleBindings are intact and
silently restores them if deleted
Fixes P1-A from integration test 2026-05-18: SA deleted from Active NS
was not being restored because reconciler only processed Failed NS.
Problem:
executor starts → AdoptExistingResources + CleanupOldExecutorObjects run
against utils.DefaultNSResolver().Snapshot() which returns ONLY static NS
from FISSION_RESOURCE_NAMESPACES. Managed (labeled) namespaces are
registered later, asynchronously, by StartNSWatcher.
Result:
- Pods from a previous executor in managed NS are never adopted
(no instanceID patch) → poolmgr creates new pool pods → cold start
for first request after executor restart.
- Old executor objects (RS/deployments) in managed NS accumulate
without being cleaned up (resource leak).
Fix:
Add multitenant.PreRegisterManagedNamespaces(ctx, logger, kubernetesClient)
called synchronously in executor.go BEFORE the adopt/cleanup goroutines.
The function does a single Namespaces.List with label
fission.io/managed=true and calls DefaultNSResolver().AddNamespace() for
each result. This is idempotent with the later watcher AddFunc calls.
Failure is non-fatal: a warning is logged and startup proceeds with
static NS only (safe degraded mode).
After this call DefaultNSResolver().Snapshot() includes managed NS, so:
- AdoptExistingResources patches old pods in managed NS with new instanceID
- CleanupOldExecutorObjects removes stale objects from managed NS
- GetReaperNamespace() returns the full tenant NS set
Files:
pkg/executor/multitenant/ns_watcher.go — PreRegisterManagedNamespaces()
pkg/executor/executor.go — call before adopt/cleanup
Problem
-------
The namespace reconciler (RunReconciler, added previously) retries namespaces
in NamespacePhaseFailed every 30s by calling DispatchResync. But the phase
could never actually reach NamespacePhaseFailed for the executor component
because the executor's NamespaceSubscriber always returned nil — swallowing
any SA-provisioning or informer-init errors. The reconciler was dead code for
the executor path.
Root cause chain
----------------
1. setupSAAndRoleBindings() — void, errors only logged internally.
2. EnsureNamespaceSA() — void, just called setupSAAndRoleBindings.
3. registerNamespace() — void, errors from both functions lost.
4. Executor AddFunc/ResyncFunc — always returned nil to dispatch().
5. dispatch() marks parts Active unconditionally → NamespacePhaseFailed
is never triggered for executor → RunReconciler never fires for executor.
Consequence: if EnsureNamespaceSA failed (transient k8s 503, RBAC webhook
timeout, etc.) the namespace appeared Active in the manager but the fetcher
ServiceAccount was missing. Pool pods would CrashLoopBackOff on every call
to that namespace until a full process restart.
Changes
-------
pkg/utils/serviceaccount.go
- setupSAAndRoleBindings: void → error. Returns the first k8s API error
so callers can decide whether to retry.
- runSACheck: ignores the error with _ = (same behaviour as before, it's
a periodic background loop that already logs internally).
- EnsureNamespaceSA: void → error, propagates setupSAAndRoleBindings.
Updated godoc to explain the retry contract.
pkg/executor/multitenant/ns_watcher.go
- registerNamespace: void → error.
* EnsureNamespaceSA error → wrapped as 'EnsureNamespaceSA: ...' and returned.
* registerExecutorTypes error → wrapped as 'registerExecutorTypes: ...' and returned.
* Success log line only emitted when both succeed.
- Added 'fmt' import for error wrapping.
pkg/executor/multitenant/namespace_subscriber.go
- AddFunc: return registerNamespace(...) instead of ignoring its error.
- ResyncFunc: same — plus a comment explaining why it is safe to call
registerNamespace again (SA creation is idempotent, executor-type
AddNamespace guards against duplicate informer creation).
pkg/utils/namespace_manager.go
- RunReconciler interface signature: added *zap.Logger parameter.
Callers pass the component logger so retries are visible in prod logs.
- RunReconciler implementation:
* Accepts logger; falls back to zap.NewNop() if nil.
* Skips the tick entirely when no failed namespaces are found (no log spam).
* Logs 'retrying failed namespaces' with count + list when found.
* Logs per-namespace 'dispatching resync'.
* Logs 'resync succeeded' or 'resync still failing, will retry' with error.
- RunManagedNamespaceWatcher: passes logger to RunReconciler.
End-to-end flow after this fix
-------------------------------
1. EnsureNamespaceSA fails (k8s 503).
2. registerNamespace returns error.
3. Executor AddFunc returns error.
4. dispatch() calls MarkPartFailed("executor") → deriveNamespacePhase →
NamespacePhaseFailed.
5. RunReconciler tick (30s) finds the namespace → DispatchResync →
registerNamespace called again → EnsureNamespaceSA (idempotent) →
if API recovered: success → MarkPartActive → NamespacePhaseActive.
6. Log line 'namespace reconciler: resync succeeded' confirms recovery.
Backward compatibility
----------------------
- NamespaceManager interface: RunReconciler gained a *zap.Logger param.
There is exactly one implementation (inMemoryNamespaceManager) and one
call site (RunManagedNamespaceWatcher). No external mocks.
- EnsureNamespaceSA: callers outside this codebase (if any) that ignore
the error will still compile (Go allows ignoring return values).
- All 26 affected tests pass: go test ./pkg/utils/... ./pkg/executor/...
./pkg/buildermgr/... ./pkg/router/...
- Add RemoveNamespace(ctx, ns) to executortype.ExecutorType interface
- Implement RemoveNamespace in poolmgr, newdeploy, container executor types
- Add per-namespace context cancellation (nsCancels map) in all three types so
informer factories are stopped when namespace is removed (fixes goroutine leak)
- Add PoolPodController.RemoveNamespace to clear envLister/podLister maps
- Add deregisterNamespace() in executor multitenant subscriber
- Switch executor/router/buildermgr watcher strategy from TrackOnly to DispatchRemove
so RemoveFunc is called when fission.io/managed label is removed
- Add RemoveFunc to executor/router/buildermgr namespace subscribers
- Add RemoveNamespace to environmentWatcher and packageWatcher with per-NS cancel
- Add RemoveNamespace to HTTPTriggerSet: cancels informers, removes from maps, calls syncTriggers
- Fix ns_watcher_test.go fakeExecutorType to implement new RemoveNamespace method
Fixes:
- Executor dedup gap: re-added namespace was silently skipped (envLister/deplLister still present)
- Goroutine/FD leak: old informer factories ran forever after namespace removal
- Router stale routes: HTTPTriggers for removed namespace stayed in routing table
- DefaultNSResolver.RemoveNamespace(): removes NS from global map on label removal
so Snapshot() and idleObjectReaper stop iterating deleted namespaces.
Fixes class of dirty-state bugs when NS name is reused by new tenant.
- HandleWatcherNamespaceRemoval: call RemoveNamespace on both TrackOnly and
DispatchRemove strategies — global resolver cleanup is always required.
- dispatch(): parallel subscriber execution via goroutine per subscriber +
sync.WaitGroup. Reduces onboarding latency from O(N_subscribers × API_latency)
to O(max(API_latency)). Safe: MarkPart* are internally mutex-protected.
- inMemoryNamespaceManager.RunReconciler(): 30s ticker scans for
NamespacePhaseFailed records and retries via DispatchResync. Started
automatically by RunManagedNamespaceWatcher. Fixes permanent stuck-failed
state caused by transient k8s API errors.
Analysis source: FORENSIC_ARCHITECTURE_AUDIT.md §Deep Risk Analysis
TestStartManagedNamespaceWatcherIntegration проверяет полный маршрут
горячей регистрации namespace без real cluster:
1. RunManagedNamespaceWatcher запускается с k8sfake.NewSimpleClientset()
2. В fake client создаётся Namespace с label fission.io/managed=true
3. Kubernetes informer детектирует событие (без polling, через Watch)
4. SubscriberFuncs.OnNamespaceAdd вызывается
5. NamespaceManager содержит запись со статусом Active
Тест доказывает, что вся цепочка
fake k8s event → informer → AddFunc → subscriber → manager
работает корректно без rolling restart процесса.
Также добавлен import metav1 в test file (требовался для CreateOptions).