- Определена архитектура managed IoT service - Согласовано решение: RabbitMQ (MVP), потом Kafka; EMQX + CRD контроллер - Создан подробный план для Sonnet (doc/iot-mvp-plan.md) - Добавлено правило в copilot-instructions: лог мышления в doc/thinking/ по датам - Полный ход рассуждений в doc/thinking/2026-04-04.md Ключевые решения: - IoT код в iot/ (легко вынести потом) - CRD IoTDevice + контроллер (как все остальное в sless) - EMQX HTTP Auth Backend для динамической аутентификации устройств - Архитектура broker-agnostic (легко переключить на Kafka) - Terraform: расширяем текущий provider (sless_iot_device ресурс)
7.2 KiB
ERR-PG-08: Concurrent Operations Not Supported
Problem Statement
After completing an Update operation on nubes_postgres resource (e.g., vault_secrets refresh), attempting to create any dependent resource immediately fails with API 422 error:
Error: Ошибка клиента
ошибка API 422: {
"DETAIL": "There is a started operation on this instance",
"TYPE": "about:blank",
"TITLE": "Concurrent operations are not supported (job status: SUCCESS)"
}
Even though WaitForOperation() has returned successfully and the previous operation is marked as completed.
Reproduction
- Terraform plan detects vault_secrets has changed (external drift)
- Execute update:
nubes_postgres.pg_test_instance: Modifying... - Update completes:
nubes_postgres.pg_test_instance: Modifications complete after 0s - Immediately try to create:
nubes_postgres_database.pg_test_db: Creating... - FAIL: API returns 422 "Concurrent operations are not supported"
Why It Happens
The Nubes API has an internal operation lock per instance. Even though the Update operation's IsSuccessful flag is true and operations are no longer in-progress, the API server is still processing asynchronous side effects:
- Vault credential updates
- Instance state synchronization
- Backend resource reconciliation
When the next Create operation is submitted, the lock is still held, causing the 422 error.
Current Symptoms
- Affects: Any sequence where Update → Create operations happen on same instance
- Timing: Happens even with 120+ second waits
- Scope: Affects DB creation, user creation, any operation on PostgreSQL instance
- Test environment: Confirmed on
k8s-3-sandbox-nubes-rurealm - Production: Unknown (not tested)
Workarounds (Current)
Workaround 1: Split Resources Across Apply Cycles
Comment out dependent resource creation, apply first update, then uncomment and apply again:
# postgres.tf
# temporarily comment out nubes_postgres_database block
# terraform apply ← creates pg_test_instance + users
# uncomment nubes_postgres_database block
# terraform apply ← creates pg_test_db
Already implemented in this project (see postgres.tf lines 87-99).
Workaround 2: Add Explicit depends_on + relies on Terraform serialization
resource "nubes_postgres_database" "pg_test_db" {
postgres_id = nubes_postgres.pg_test_instance.id
db_name = var.pg_db_name
db_owner = nubes_postgres_user.pg_test_user.username
# Explicit depends_on forces sequential execution
# but does NOT help with concurrent operation lock
depends_on = [nubes_postgres_user.pg_test_user3]
}
Status: Doesn't solve the problem - API still returns 422.
Workaround 3: Manual Sequential Runs
# First apply - creates instance and users
terraform apply -auto-approve
# Wait manually (or check instance state)
sleep 180
# Second apply - creates databases
terraform apply -auto-approve
This is unreliable and not automatable.
Root Cause Analysis
Client-Side (Terraform Provider)
File: internal/provider/client_impl.go line 240
func (c *NubesClient) WaitForOperation(ctx context.Context, opUid string) error {
timeout := time.After(15 * time.Minute)
ticker := time.NewTicker(10 * time.Second)
// ...
// Checks every 10 seconds for operation completion
if !op.IsInProgress && !op.IsPending && op.DtFinish != nil {
if op.IsSuccessful != nil && *op.IsSuccessful {
return nil // ← Returns immediately when successful
}
}
}
Problem: No post-completion delay or retry logic for subsequent operations.
Server-Side (Nubes API)
The API maintains an operation lock on the instance that:
- Is released when operation completes (
IsSuccessful = true) - But async background tasks are still running during the lock release window
- New requests during this window: "There is a started operation on this instance"
This is an intentional safety measure against corrupting instance state, but the window between "operation done" and "instance ready for next operation" is not deterministic.
Solutions (For Provider Fix)
Option 1: Add Post-Completion Delay
Pros: Simple, guaranteed to work Cons: Always adds overhead, even if not needed
// In client_impl.go, after "return nil" on success:
if op.IsSuccessful != nil && *op.IsSuccessful {
// Add buffer for API server to release internal locks
time.Sleep(30 * time.Second) // or configurable
return nil
}
Recommended value: 30-60 seconds based on observations.
Option 2: Implement Retry Mechanism
Pros: No unnecessary delays, adapts to actual API response time Cons: More complex, needs careful timeout/backoff tuning
When next operation fails with "concurrent operations", retry with exponential backoff:
func (c *NubesClient) CreateResourceWithRetry(ctx context.Context, payload map[string]interface{}) error {
maxRetries := 5
backoff := 10 * time.Second
for i := 0; i < maxRetries; i++ {
err := c.Create(ctx, payload)
if err == nil {
return nil
}
if strings.Contains(err.Error(), "Concurrent operations") {
time.Sleep(backoff)
backoff *= 2 // exponential backoff
continue
}
return err
}
}
Backoff suggestion: Start 10s, cap at 60s.
Option 3: Query Instance State Before Next Operation
Pros: Most elegant, confirms instance is ready Cons: Requires additional API call, might still be unreliable
func (c *NubesClient) WaitForInstanceReady(ctx context.Context, instanceId string) error {
// Poll instance state directly, not just operation state
for retry := 0; retry < 30; retry++ {
state, err := c.GetInstanceState(ctx, instanceId)
if err == nil && state.IsReady {
return nil
}
time.Sleep(5 * time.Second)
}
return fmt.Errorf("instance not ready after timeout")
}
Recommendation
Implement Option 1 (Post-Completion Delay) combined with Option 2 (Retry Logic):
- Add fixed 30-second delay after
WaitForOperationreturns success (fail-safe) - Keep retry mechanism for cases where clients don't respect the delay (defensive)
This provides both reliability (fixed delay) and robustness (retry on failure).
Testing
Test case:
cd examples/PG_TEST
# Uncomment pg_test_db in postgres.tf
terraform apply -auto-approve
# Should NOT fail with 422 "Concurrent operations are not supported"
# Should create all 3 resources: instance, users, database
Current status: ❌ FAILS with 422
After fix: ✅ SHOULD PASS
References
- Provider source:
/home/naeel/terra/terraform/internal/provider/ - Test configuration:
/home/naeel/terra/sless/examples/PG_TEST/ - Related: ERR-PG-02 (fixed), ERR-PG-03 (race condition), ERR-PG-04 (invalid role)
- Vault credentials: Not involved in this error (different subsystem)
Date discovered: 2026-04-03
Status: Open, blocker for multi-resource deployments
Priority: High (blocks full lifecycle automation)
Scope: Test environment confirmed, production unknown