Architecture & product experiment

Threadz

Email intelligence for small businesses

The hard part was not making the AI answer.
It was deciding when the system had earned the right to answer.

Active prototype / field validation

Small businesses often run on email. Threadz asks whether that messy history can become useful operational information without allowing AI to guess.

Email is good at communication.
Less good at operational memory.

Client requests, invoice details, commitments, follow-ups and decisions can be spread across long threads, attachments and individual inboxes. A business owner may remember the conversation but still struggle to find the exact detail.

CRM and finance systems help once information has been entered correctly. Threadz explores whether useful structure can be recovered from the email itself, while preserving evidence and uncertainty.

  • Important information buried in threads and attachments
  • Inconsistent client names and references spread across messages
  • Commitments expressed in everyday language, with dates that mean different things
  • Plausible AI answers that are not actually supported by the source

The first idea was simple

Extract useful facts from email and present them in a simple business workspace. The obvious starting point was to ask an AI model to do the extraction.

It worked just well enough to be dangerous. Answers could look convincing even when testing showed they were wrong. In the local comparisons, a stronger model did not reliably remove semantic mistakes. Stricter trust rules blocked unsafe claims but also rejected useful results.

Adding conversational context did not consistently fix the problem. The architecture changed direction: make the task smaller, and validate each decision separately.

Make the problem smaller

The experiment separated deterministic parsing from narrowly scoped AI extraction. Deterministic means applying explicit rules to information the source actually contains. AI candidates still have to earn their place in a record.

Sender and recipient metadata are parsed directly. Invoice and PO labels are read structurally. Amounts and currencies are separate fields. A message timestamp is not treated as the date of a business event. Evidence is attached to each surfaced value, and uncertain values can remain unknown.

If a fact cannot be supported by evidence, do not promote it into the business record.
Threadz evidence architecture
The broader experiment includes action and date research. The current client interface exposes only the retained invoice slice. Open full resolution ↗

What 13 iterations taught us

Reliability improved when the system became less ambitious. These are findings from bounded development samples, not a general benchmark for all email or every model.

  • Larger local models were not an automatic route to reliable facts; unsupported claims and commitment/completion mistakes remained.
  • Tighter trust policies traded useful coverage for fewer unsafe answers. Evidence from metadata needed to be distinguished from evidence in body text.
  • Free-form clarification and broader context did not consistently resolve the underlying understanding problem.
  • Specialist extraction and structural label/value parsing improved invoice-reference recovery.
  • Explicit label precedence stopped competing PO labels from being promoted as invoice references in the reviewed cases.
  • Safe abstention was preferable to forcing a link, but missed references and excessive abstention remained visible failures.

A trustworthy system needs constraints on what the model may decide, not just more confident wording. The labels were agent-reviewed rather than independently adjudicated by the client; reused examples and concentrated formats limit how far these results generalise.

What Threadz can safely do today

The limited invoice slice met its selected-sample gates for surfaced records. It can:

  • Detect invoice identifiers with high precision in the reviewed sample, while still missing some references
  • Show supported GBP amounts and currencies against a particular invoice mention
  • Link attachment metadata by filename evidence; attachment contents are not read or verified
  • Associate amounts and linked filenames with an invoice where the retained rules support that connection
  • Link PO references conservatively; coverage remains partial
  • Keep exact source evidence behind each surfaced value
Records remain separate per source email, not globally reconciled invoices. Client matching remains limited. No reliable payment status, overdue detection, action completion, general business-state reasoning or autonomous follow-up is claimed. Missing receipts do not establish non-payment.

The current prototype deliberately omits capabilities that have not passed their reliability gates.

From prototype to production

Threadz is an architecture and product experiment, not a production software-engineering exercise. The work focuses on shaping the problem, defining system boundaries, using AI-assisted prototyping to test assumptions, and deciding what the evidence actually supports.

The prototype explores which decisions should be deterministic, where AI adds value, when it should abstain, how evidence should be retained and where human judgement belongs.

A production implementation would require a specialist engineering team to harden the system around that architecture. The diagram below is a possible design, not a set of production capabilities implemented or operated by the lab.

Proposed Threadz production architecture
Illustrative engineering path. Provider integrations, tenant isolation and production operations have not been demonstrated by the local prototype. Open full resolution ↗

What production engineering would add

  • Secure OAuth and token lifecycle management, scoped provider access and secrets management
  • Tenant isolation, authorisation and encryption at rest and in transit
  • CI/CD with unit, integration and security tests; model and version evaluation gates
  • Observability, privacy-aware logs and an auditable record of processing decisions
  • Resilience, retry behaviour, duplicate handling, backups and recovery
  • GDPR and data-retention controls, deletion processes, threat modelling and security review
  • Scale and performance testing, with a defined production support model

These are engineering requirements to design, implement and verify, not production credentials claimed by this experiment.

Agentic AI should not be defined by how much autonomy the system has.
It should be defined by how deliberately autonomy is bounded.Deterministic where possible. AI where useful. Human judgement where necessary.

Controlled real-user validation

The next gate is whether the bounded invoice slice helps a small business owner recover useful context more quickly. Representative email evidence, and independently reviewed CRM or invoice records where available, would provide a reference for that assessment. No CRM integration is implied.

Potential directions include stronger client matching, invoice reconciliation, explicit payment-state extraction, safer action tracking, CRM integration and Microsoft 365 or Gmail connectors. These are areas to investigate, not committed features.

Further capabilities will only be introduced when they can demonstrate acceptable reliability.

Back to Labs

Try the demonstrator

Explore the product concept with a separate, entirely fictional workspace.

About the Threadz Build
Launch Threadz demo (opens in a new tab) ↗