Middleware works as an intermediary runtime: it listens for triggers, fetches or receives payloads, applies technical and business rules, then delivers results to one or more target systems. The value is reliability and clarity—especially when formats differ or multiple steps must succeed in order.
Think of middleware as air-traffic control for data, not as a second database of record (unless you deliberately design caching). For when a hub is justified, see what is middleware. For hub vs direct links, see middleware vs direct API integration.
Trigger models
- Schedule-based — pull new or changed records every N minutes
- Event-based — react to webhooks or queue messages
- On-demand — run when a user clicks “sync” or an admin job starts
Mixed models are common: webhooks for speed, scheduled reconciliation for safety.
Inside a processing pipeline
Inside a middleware pipeline
- Ingest event or poll for changes
- Authenticate and identify tenant or company
- Validate schema and required fields
- Transform, map, and enrich the payload
- Call target APIs with idempotent keys
- Log outcome; retry, quarantine, or alert
Authentication and tenancy
Each hop needs correct credentials—API keys, OAuth tokens, certificates—scoped to least privilege. If one middleware instance serves multiple companies or branches, isolate credentials, configuration, and logs per tenant. Cross-tenant leakage is an existential failure.
Transform and map at the edges
Prefer a clear internal shape for a business event (“SalesOrderConfirmed”) and translate to each vendor’s API at the edges. Canonical models reduce the cost of swapping one ecommerce platform later. Keep data mapping in config where possible so finance can review field rules without reading code.
Retries for transient failures
Networks fail. APIs time out. Retry with backoff for timeouts and rate limits; do not blindly retry permanent validation errors. Cap retry windows and escalate when automatic recovery cannot finish inside the SLA.
Idempotency
Assume duplicates and partial failures. Use idempotency keys, upserts, and natural business keys (invoice number + supplier) so a replay does not double-post. Without idempotency, “safe retry” becomes a finance incident.
Quarantine (dead-letter) for poison messages
Bad data that will never succeed should not block the entire queue. Quarantine poison messages with enough context—payload snapshot, error, correlation ID—for a human to fix mapping or source data, then replay deliberately.
Logging and observability
Operators need to answer: What ran? What failed? Which record? Can we replay? Use correlation IDs across hops. Log enough to debug; avoid retaining full sensitive payloads forever without policy. Dashboards for success/failure counts, lag, and oldest unprocessed event keep trust high.
Replay tools
Failed or quarantined messages need a controlled replay path—single message, filtered batch, or time window—after the root cause is fixed. Silent “run the whole sync again” without idempotency is how duplicates appear.
Multi-step orchestration
Some flows require ordered calls (create customer, then invoice, then payment allocation). Middleware should track step state, compensate or alert on partial success, and never leave operators guessing which step completed.
FAQ
How do we know middleware is healthy?
Dashboards for success/failure, lag, and oldest unprocessed event—plus synthetic canaries that post a known test document on a schedule.
Should middleware store business data permanently?
Usually no. Cache or staging is fine; the systems of record should remain accounting, CRM, or ecommerce. Long-term stores need an explicit retention reason.
What is the difference between retry and replay?
Retry is automatic for transient errors. Replay is a deliberate operator action after fixing data or mapping.
How do webhooks and schedules work together?
Webhooks reduce latency for new events; schedules catch missed deliveries and reconcile drift. Many production hubs use both.
Where do compliance submission APIs fit?
As target (or source) connectors with the same auth, idempotency, and logging rules. For a portal-vs-API discussion, see MyInvois portal vs API integration & middleware.
Related concepts
Definition and when to use a hub: what is middleware. Pattern choice: middleware vs direct API. Sync concepts: what is data synchronization.