How We Reduced Email Processing Latency by 10x
The Starting Point
When SpiderMail launched, our email processing pipeline took an average of 200ms per message. For a single email, that's fine. For an agent processing thousands of messages per hour, it's a bottleneck.
Profiling the Pipeline
We instrumented every stage:
MIME parsing: 45ms
HTML sanitization: 78ms
Content extraction: 32ms
YAML serialization: 15ms
Metadata enrichment: 30ms
HTML sanitization was the clear bottleneck.
The Optimizations
1. Streaming Parser
We replaced our DOM-based HTML parser with a streaming SAX parser. Instead of building a full DOM tree and then walking it, we process tokens as they arrive.
2. Compiled Sanitization Rules
Our 2,000+ sanitization rules were being interpreted at runtime. We now compile them into a finite-state machine at startup.
3. Zero-Copy YAML Output
Instead of building strings and concatenating them, we write directly to a pre-allocated buffer.
4. Parallel Metadata Enrichment
DNS lookups, reputation checks, and header analysis now run concurrently.
Results
| Stage | Before | After | |-------|--------|-------| | MIME parsing | 45ms | 8ms | | HTML sanitization | 78ms | 4ms | | Content extraction | 32ms | 3ms | | YAML serialization | 15ms | 1ms | | Metadata enrichment | 30ms | 2ms | | Total | 200ms | 18ms |
The 10x improvement means our customers can process 200,000+ emails per hour on a single node.