Skip to content
VZU
VZU Custom Software & Cloud · Pillar 05

Data Engineering.Real-time, not yesterday.

Real-time ETL. Streaming pipelines. Data lakehouse. Decisioning. Anomaly detection. Unified BI. 70% faster reporting cycles. The brief is the contract. The work is the work.

Pillar

05 · Data Engineering

70%

Faster reporting

How Data Engineering runs on the VZU runtime

KAI builds the pipeline. QUILL reconciles the data.

The Data Engineering pillar is operated by two agents from the VZU runtime. KAI is the backend and infrastructure engineer. QUILL is the spreadsheet and data reconciliation engineer. Together they ship a real-time, lakehouse-grade data platform with the audit trail on from the first row.

Agent 01 · VZU Stack

KAI

Backend + infrastructure engineer. Database, API, queue, worker, frontend, infra, deploy, monitoring. The agent you call when the work is foundational.

  • mcp.db.migrate — apply database migrations
  • mcp.api.deploy — deploy the API to staging/prod
  • mcp.queue.worker — start the background worker
  • mcp.infra.terraform — infrastructure as code
  • mcp.monitor.grafana — monitoring dashboards

Median brief

10 weeks

Services deployed

47

Uptime

99.97%

Migrations

12

Agent 02 · VZU Sheets

QUILL

Spreadsheet engineer. Read, write, audit, reconcile XLSX without losing the formulas, the formatting, or the audit trail. The agent that turns spreadsheets into APIs.

  • mcp.xlsx.read — XLSX preserving formulas + formatting
  • mcp.xlsx.write — XLSX preserving structure
  • mcp.xlsx.reconcile — reconcile two XLSX files
  • mcp.xlsx.audit — audit XLSX for errors

Median brief

2 weeks

Sheets processed

23

Rows reconciled

14,000

Reconciliation rate

99.7%

Orchestration

KAI → stand up lakehouse  ·  KAI → wire streaming pipelines  ·  QUILL → reconcile source + target  ·  KAI → deploy BI  ·  QUILL → audit final data. Brief completed.

What Data Engineering ships

Eight things this practice does, end to end.

VZU Data Engineering offerings
Real-time ETL

Streaming, not batch

Kafka, Pulsar, Kinesis, Pub/Sub. Sub-second latency from event to warehouse. The dashboard reflects the shelf, not yesterday's shelf.

Data lakehouse

Iceberg, Delta, Hudi

Open table formats on S3, GCS, ADLS. ACID transactions. Time travel. Schema evolution. The lake is the warehouse, and the warehouse is the lake.

Streaming pipelines

Flink, Spark Structured Streaming

Exactly-once semantics. State management. Windowed aggregations. The pipeline is the brief, made concrete. The agent runs it 24/7.

Decisioning

AI-driven at the SKU, the user, the moment

Recommendations, fraud scoring, churn prediction, dynamic pricing. Models deployed at the edge, scoring sub-50ms. The decision is the data, in motion.

BI & dashboarding

Looker, Tableau, Metabase, custom

Self-serve for the operator. Governed for the enterprise. The dashboard is a question, answered. The data warehouse is the source of truth, not the report.

Anomaly detection

AI-driven, severity-ranked

Statistical, ML, and LLM-based. The agent watches every metric, surfaces the anomaly, ranks it by business impact, and writes the remediation plan.

Data governance

Lineage, catalog, RBAC, audit

OpenLineage, DataHub, Unity Catalog, custom. The data lineage is the audit trail. The catalog is the source of truth. The RBAC is the operator's, not the vendor's.

Migration & modernization

Legacy warehouse → lakehouse

Teradata, Netezza, Oracle Exadata, SQL Server. Migrated with zero downtime, dual-write validation, and the audit trail on from the first row.

VZU Data Engineering use case — a Fortune 500 retailer at 70% faster reporting
Production

Use case · F500 Retailer

70% faster reporting. Real-time decisioning at the shelf.

A Fortune 500 retailer had a 36-hour reporting cycle and a 4-hour daily close. Decisions were made on stale data. The VZU Data Engineering practice rebuilt the pipeline end to end. Real-time ETL with Kafka and Flink. Lakehouse on Iceberg + Snowflake. Streaming decisioning at the SKU level. AI-driven anomaly detection that surfaces the issue before the manager notices. Unified BI on Looker. Reporting cycles dropped from 36 hours to under 11. The daily close went from 4 hours to 47 minutes. 12M+ SKUs are now tracked in real time. The shelf is the dashboard, and the dashboard is the shelf.

  • Real-time ETL, streaming data lakehouse, unified BI
  • AI-driven anomaly detection at the SKU level
  • 36-hour reporting cycle → under 11 hours
  • 70%

    Faster reporting

  • 47min

    Daily close

  • 11hrs

    Reporting cycle

  • 12M+

    SKUs tracked

Brief VZU on Data Engineering