Build modular data pipelines with the clarity of Unix pipes

Introducing an innovative open-source prototype that applies Unix philosophy to machine learning pipelines, we have designed a modular system where each stage—such as PII redaction, chunking, deduplication, embeddings,…

3 min readMachine Learning

The Unix pipe is elegant because it does one thing: it lets you chain small, reliable tools without worrying about what happens inside them. That same clarity has been missing from retrieval pipelines, where swapping a single component often means untangling a web of hidden dependencies. The prototype from this open-source project tackles that problem directly, and we think it points in the right direction.

The team's motivation is something any practitioner will recognize. They swapped a chunker, saw retrieval quality drop, and could not tell whether the chunker itself was at fault or whether it had broken something downstream. That ambiguity is expensive. It turns debugging into guesswork and slows iteration to a crawl. By treating each stage, PII redaction, chunking, dedup, embeddings, evaluation, as an independent plugin with a typed contract, this design forces clarity. You change one option, re-run the eval, and compare precision and recall directly. No more wondering whether the embedding model silently shifted behavior because the chunk boundaries changed.

What makes this approach practical is the explicit stage boundary in the feature naming convention. `docs__pii_redacted__chunked__deduped__embedded__evaluated` is not just a label; it is a contract. Each double underscore marks a point where data must conform to a known shape. Swap the chunking method from `sentence` to `paragraph`, and the rest of the pipeline stays exactly as it was. The evaluation results tell you what that single change did, nothing else. That is the Unix philosophy applied to data pipelines: small, composable pieces with well-defined interfaces.

We should be clear that this is still a prototype, not a production system. The team is asking for feedback on whether the design assumptions hold up, and that is the right posture. The value here is not in a finished product but in a principle: modularity enforced by typed contracts makes retrieval pipelines testable, debuggable, and improvable. If the community can validate that this approach scales beyond a prototype, it will change how we think about building RAG systems. For now, the question worth answering is whether your own pipeline can survive a single component swap without breaking something you did not expect.

From Machine Learning

We built an open-source prototype that applies Unix philosophy to retrieval pipelines. Each stage (PII redaction, chunking, dedup, embeddings, eval) is its own plugin with a typed contract, like pipes between Unix tools. The motivation: we swapped a chunker and retrieval got worse, but could not isolate whether it was the chunking or something breaking downstream. With each stage independently swappable, you change one option, re-run eval, and compare precision/recall directly. ```python Feature("docs__pii_redacted__chunked__deduped__embedded__evaluated", options={ "redaction_method": "presidio", "chunking_method": "sentence", "embedding_method": "tfidf", }) ``` Each `__` is a stage boundary. Swap any piece, the rest stays the same. Still a prototype, not…

Read the original at Machine Learning