Data & interfaces
The weekly report stopped eating someone's Friday
Data from email, spreadsheets and CRM collected automatically, mapped to one metric dictionary, and opened as a dashboard you can ask "why".
- Task
- Automated collection and parsing of operational data
- Timeline
- 5 working days
- Stack
- n8n · OpenAI · TypeScript · PostgreSQL
- Result
- 9.2 s per document · 3.8 s per answer
The problem
Every week, several people manually move data from emails, attachments and spreadsheets into one report. It is ready by Friday evening, stale by Monday, and different departments calculate the same metric in different ways.
External benchmarks:
- A 2025 survey reports that employees spend more than nine hours per week moving data from emails, PDFs, spreadsheets and scans into digital systems; the total annual cost was estimated at $28,500 per employee. These are external survey figures.
- In finance, the figure is higher: 41% of FP&A processes are manual, or about 10 hours per week of skilled time in spreadsheets. These are external survey figures.
What was built
- Scheduled collection — n8n pulls data from email, spreadsheets and the CRM, normalizes it and stores it in PostgreSQL.
- Unstructured attachment parsing — an AI agent turns an email, PDF or export into fields that follow a defined schema. Measured: 9.2 seconds per document in a single run, HTTP 200.
- Metric dictionary — each metric has one definition, one source and a record of who agreed it and when. Every number on the dashboard shows where it came from.
- TypeScript dashboard — metrics, trends and business-unit views update on schedule rather than on a button press.
- AI analyst beside the numbers — a natural-language question returns an analysis with hypotheses. Measured: 3.8 seconds per answer in a single run.
- Required interface note — AI hypotheses must be checked against source data. This is a condition for trusting the dashboard, not decoration.
What happens when something breaks
- A source fails to provide data — the metric is explicitly marked “incomplete data for this period” instead of being silently calculated from a partial set. A quietly wrong number is worse than a missing one.
- Document parsing fails — the document moves to a manual-review queue with a link to the original.
- The dashboard differs from the source system — a scheduled reconciliation triggers an alert above the agreed threshold.
- The analyst is unavailable — the dashboard keeps working because its metrics do not depend on model availability.
Results
| Metric | Before | After | Nature of the figure |
|---|---|---|---|
| Manual report assembly | ~12 h/week for the team | ~1 h/week for checking | calculated model |
| Data freshness | Weekly | Scheduled, daily | architecture property |
| Incoming document parsing | 5–10 min manually | 9.2 s | measured |
| Answer to a metric question | “Ask the analyst; they will check” | 3.8 s | measured |
| Definitions of one metric | 2–3 versions | One, with a named source | architecture property |
How this was calculated
Before: 3 people × 4 h/week building the report = 12 h/week ≈ 48 h/month
After: 1 h/week checking and commenting ≈ 4 h/month
Time returned ≈ 44 h/month
Australian analyst cost ≈ AUD 55/h fully loaded (≈ USD 36/h)
Monthly value ≈ 44 × USD 36 ≈ USD 1,584/month (≈ AUD 2,420)
Build payback = USD 1,190 ÷ USD 1,584 ≈ 0.8 months
With maintenance = USD 1,190 ÷ USD 1,434 ≈ 0.8 months
Assumptions: three people, four hours a week each, an operations analyst at AUD 55/h including super and on-costs, 1 AUD = 0.65 USD. The external benchmark of more than nine hours per week on manual data transfer shows that four hours is a conservative estimate, rather than an inflated one.
What did not work on the first attempt
- The first version calculated “conversion” from two systems with different definitions of a lead. The numbers did not match the CRM, and trust in the dashboard disappeared in one day. The fix is a metric dictionary and a visible source for every number, rather than adjusting the formula to match.
- The AI analyst eagerly explained random variation as a trend. The instruction now requires sample size, and the interface warns that hypotheses need verification.
- Daily refresh was excessive for some metrics and created alert noise. Refresh frequency became a metric-level setting rather than a system-wide setting.
Timeline and cost
| Day | Work | Deliverable |
|---|---|---|
| 1 | Inventory sources, define the metric dictionary and acceptance | Agreed metric definitions |
| 2 | Collection, normalization and storage | Data in the database, reconciled with sources |
| 3 | Parse unstructured documents | Agent on real files |
| 4 | Dashboard, views, analyst and alerts | Working dashboard |
| 5 | Reconcile with source systems, document and record a walkthrough | System and metric dictionary in the owner’s account |
External validation
- Survey: manual data entry costs US companies $28,500 per employee each year — more than nine hours per week moving data.
- DataRails: 41% of FP&A processes are manual — about 10 hours per week in spreadsheets.
- MIT Project NANDA, The GenAI Divide (2025) — why projects without process instrumentation fail to produce measurable effect.
Try it
Do you have a similar process?