This is a review for backend engineers who already know Go and need to discuss or build production services again. The useful decisions are bounded concurrency, cancellation, finite queues, idempotency, and observability.
The scope is Go 1.26. It is not a language introduction. Move quickly through syntax and spend time on the choices that shape a consumer, API, or telemetry pipeline.
30-minute route
| Time | Topic | Priority |
|---|---|---|
| 0-4 min | Types, structs, interfaces, errors | Quick review |
| 4-10 min | Goroutines, channels, select, context | High |
| 10-17 min | Bounded concurrency and backpressure | Highest |
| 17-23 min | Resilient IoT pipeline | Highest |
| 23-26 min | Runtime, memory, profiling | High |
| 26-30 min | Architecture and interview answers | Highest |
Production fundamentals
Go favors composition, small contracts, and explicit flow. A value can validate itself without a framework:
var ErrOutOfRange = errors.New("reading out of range")
type Reading struct {
DeviceID string
Sequence uint64
Value float64
}
func (r Reading) Validate() error {
if r.DeviceID == "" {
return errors.New("device_id is required")
}
if r.Value < -100 || r.Value > 250 {
return fmt.Errorf("%w: %.2f", ErrOutOfRange, r.Value)
}
return nil
}
The zero value is often useful. Slices share a backing array until append reallocates; maps have no iteration order and need synchronization for concurrent access. Strings are immutable bytes, usually UTF-8. defer runs in LIFO order, while evaluating its arguments when registered.
Interfaces are satisfied implicitly. Define a small interface where it is consumed. Errors are values: add context with %w and inspect the chain with errors.Is or errors.As. Reserve panic for broken invariants or unrecoverable startup failures.
Goroutines, channels, and context
A goroutine is not a dedicated OS thread. The runtime schedules it over OS threads. Every goroutine needs a clear owner, stop condition, and wait path.
Channels move work or ownership. Mutexes protect shared state. A buffered channel smooths a temporary speed difference; it does not create unlimited capacity.
func enqueue(ctx context.Context, jobs chan<- Reading, reading Reading) error {
select {
case jobs <- reading:
return nil
case <-ctx.Done():
return context.Cause(ctx)
}
}
The producer closes a channel when it knows no more values will be sent. Sending to a closed channel or closing it twice panics. A nil channel blocks forever, and disables its select case.
context.Context carries cancellation, deadlines, and request-scoped metadata. Receive it first, propagate it, call every returned cancel, and do not store it in a struct. Cancellation is cooperative: blocking loops must watch ctx.Done().
Bound concurrency before memory becomes the limit
One goroutine per message becomes expensive when a downstream slows down. Queues and heap grow, GC gets busier, and the process may fail before CPU looks full.
errgroup combines waiting, first-error propagation, and shared cancellation. Set a limit for independent tasks:
func ProcessBatch(ctx context.Context, batch []Reading) error {
g, ctx := errgroup.WithContext(ctx)
g.SetLimit(16)
for _, reading := range batch {
reading := reading
g.Go(func() error {
return processOne(ctx, reading)
})
}
return g.Wait()
}
Classify errors before acting. A database outage can cancel a batch. Invalid, duplicate, or schema-incompatible messages belong in quarantine or a DLQ, not in a failure that stops the whole consumer.
Use atomic for an independent flag or counter, sync.Mutex for an invariant across fields, and a channel for work transfer. Do not copy a mutex after first use or keep a lock during remote I/O.
Backpressure is a product and operations policy. When ingestion accepts 50,000 messages per second and persistence completes 20,000, storing the difference in memory only moves the incident. Decide whether to block producers, reject with retry, pause consumption so the broker holds durable backlog, discard stale samples, aggregate, or spill to disk. Define queue size, occupancy metric, timeout, and saturation action.
A resilient IoT pipeline
device -> MQTT/broker -> Go ingestion -> stream -> processors -> storage
\-> DLQ \-> current state
MQTT fits device connectivity. A stream such as Kafka fits durable retention, replay, and internal partitioning. gRPC is typed internal RPC; WebSocket updates dashboards. These protocols solve different boundaries.
Multiple workers break global ordering. Telemetry commonly needs order per device, so partition by a stable key such as hash(device_id) % N and process each partition sequentially. Keep both observed_at and ingested_at, plus sequence, event_id, and boot_id where applicable. Device clocks drift and restart.
Treat the end-to-end path as at least once. Receive the event, validate its envelope and schema version, check an idempotency key, persist the effect and deduplication marker in one transaction where possible, then ACK. A transactional outbox closes the gap between committing database state and publishing a following event. Consumers still need idempotency because duplicates remain possible.
Retry only transient failures. Invalid payloads and rejected business rules do not improve with another attempt. Add a limit, a total budget, and jitter so replicas do not retry together.
At the edge, use TLS, per-device identities, topic authorization, strict payload limits, and validation before allocating large structures. Rotation, revocation, sequence numbers, nonces, and time windows matter when the business protocol must resist replay.
Kubernetes, observability, and performance
On SIGTERM, remove readiness, stop fetching work, drain in-flight work within the grace period, ACK only completed messages, and close producers, connections, and telemetry last. Liveness asks whether the process progresses; it should not depend on every external service. Readiness asks whether this pod can accept work now.
For consumer autoscaling, CPU alone is weak. Watch lag, age of the oldest message, arrival rate, processing time, and worker-pool occupancy.
Use structured logs and correlation fields without logging credentials or full sensitive payloads. Track throughput, errors by class, p50/p95/p99, lag, event age, retries, DLQ volume, duplicates, goroutines, heap, and GC pauses. Use sampled traces across ingestion, stream, and persistence; tracing every high-frequency reading can cost more than it helps.
A data race is concurrent access to one memory location with at least one write and no synchronization order. Channel sends, mutex unlock/lock, and atomic operations create useful ordering. The detector only covers executed paths:
go test -race ./...
go test -bench=. -benchmem ./...
go tool pprof cpu.out
go tool trace trace.out
G is a goroutine, M an OS thread, and P a logical execution resource. GOMAXPROCS limits Ps that run Go code concurrently, not the goroutine count. Goroutines are lightweight, not free. Profile before pooling or micro-optimizing. Preallocate known slice capacity, avoid repeated string/[]byte conversions in hot paths, and treat sync.Pool as an opportunistic temporary-object cache.
Interview answers to keep ready
Is a goroutine a thread? No. It is a lightweight runtime-managed execution unit multiplexed over OS threads.
Channel or mutex? Use a channel to transfer work or ownership; use a mutex to protect shared state and invariants.
Who closes a channel? The producer that knows there will be no more sends.
How do you preserve ordering with workers? Avoid global ordering unless it is required. Partition by a key such as device_id and process each partition sequentially.
How do you handle duplicates? Use a stable idempotency key, transactional deduplication where possible, and naturally idempotent updates such as an upsert with version or sequence.
How do you investigate latency? Separate queue time, processing, and dependencies. Compare p95/p99, lag, and saturation, then test a hypothesis with tracing, pprof, or go tool trace.
Production checklist
- Does every goroutine have an owner and a stop condition?
- Where do context cancellation and deadlines propagate?
- What is the concurrency limit and what happens at saturation?
- Is ordering required per key or globally?
- When does ACK happen, and how does the mutation survive re-delivery?
- Which failures retry, go to DLQ, or go to quarantine?
- How do device time and ingestion time differ?
- Which metrics expose lag, p99, heap, GC, and contention?