Data Engineering.Real-time, not yesterday.
Real-time ETL. Streaming pipelines. Data lakehouse. Decisioning. Anomaly detection. Unified BI. 70% faster reporting cycles. The brief is the contract. The work is the work.
Pillar
05 · Data Engineering
70%
Faster reporting
How Data Engineering runs on the VZU runtime
KAI builds the pipeline. QUILL reconciles the data.
The Data Engineering pillar is operated by two agents from the VZU runtime. KAI is the backend and infrastructure engineer. QUILL is the spreadsheet and data reconciliation engineer. Together they ship a real-time, lakehouse-grade data platform with the audit trail on from the first row.
Agent 01 · VZU Stack
KAI
Backend + infrastructure engineer. Database, API, queue, worker, frontend, infra, deploy, monitoring. The agent you call when the work is foundational.
- →mcp.db.migrate — apply database migrations
- →mcp.api.deploy — deploy the API to staging/prod
- →mcp.queue.worker — start the background worker
- →mcp.infra.terraform — infrastructure as code
- →mcp.monitor.grafana — monitoring dashboards
Median brief
10 weeks
Services deployed
47
Uptime
99.97%
Migrations
12
Agent 02 · VZU Sheets
QUILL
Spreadsheet engineer. Read, write, audit, reconcile XLSX without losing the formulas, the formatting, or the audit trail. The agent that turns spreadsheets into APIs.
- →mcp.xlsx.read — XLSX preserving formulas + formatting
- →mcp.xlsx.write — XLSX preserving structure
- →mcp.xlsx.reconcile — reconcile two XLSX files
- →mcp.xlsx.audit — audit XLSX for errors
Median brief
2 weeks
Sheets processed
23
Rows reconciled
14,000
Reconciliation rate
99.7%
Orchestration
KAI → stand up lakehouse · KAI → wire streaming pipelines · QUILL → reconcile source + target · KAI → deploy BI · QUILL → audit final data. Brief completed.
What Data Engineering ships
Eight things this practice does, end to end.

- Real-time ETL
-
Streaming, not batch
Kafka, Pulsar, Kinesis, Pub/Sub. Sub-second latency from event to warehouse. The dashboard reflects the shelf, not yesterday's shelf.
- Data lakehouse
-
Iceberg, Delta, Hudi
Open table formats on S3, GCS, ADLS. ACID transactions. Time travel. Schema evolution. The lake is the warehouse, and the warehouse is the lake.
- Streaming pipelines
-
Flink, Spark Structured Streaming
Exactly-once semantics. State management. Windowed aggregations. The pipeline is the brief, made concrete. The agent runs it 24/7.
- Decisioning
-
AI-driven at the SKU, the user, the moment
Recommendations, fraud scoring, churn prediction, dynamic pricing. Models deployed at the edge, scoring sub-50ms. The decision is the data, in motion.
- BI & dashboarding
-
Looker, Tableau, Metabase, custom
Self-serve for the operator. Governed for the enterprise. The dashboard is a question, answered. The data warehouse is the source of truth, not the report.
- Anomaly detection
-
AI-driven, severity-ranked
Statistical, ML, and LLM-based. The agent watches every metric, surfaces the anomaly, ranks it by business impact, and writes the remediation plan.
- Data governance
-
Lineage, catalog, RBAC, audit
OpenLineage, DataHub, Unity Catalog, custom. The data lineage is the audit trail. The catalog is the source of truth. The RBAC is the operator's, not the vendor's.
- Migration & modernization
-
Legacy warehouse → lakehouse
Teradata, Netezza, Oracle Exadata, SQL Server. Migrated with zero downtime, dual-write validation, and the audit trail on from the first row.
Production Use case · F500 Retailer
70% faster reporting. Real-time decisioning at the shelf.
A Fortune 500 retailer had a 36-hour reporting cycle and a 4-hour daily close. Decisions were made on stale data. The VZU Data Engineering practice rebuilt the pipeline end to end. Real-time ETL with Kafka and Flink. Lakehouse on Iceberg + Snowflake. Streaming decisioning at the SKU level. AI-driven anomaly detection that surfaces the issue before the manager notices. Unified BI on Looker. Reporting cycles dropped from 36 hours to under 11. The daily close went from 4 hours to 47 minutes. 12M+ SKUs are now tracked in real time. The shelf is the dashboard, and the dashboard is the shelf.
- → Real-time ETL, streaming data lakehouse, unified BI
- → AI-driven anomaly detection at the SKU level
- → 36-hour reporting cycle → under 11 hours
-
70%
Faster reporting
-
47min
Daily close
-
11hrs
Reporting cycle
-
12M+
SKUs tracked
Adjacent pillars