neesh Inc.
Scenarios

Your property management inbox handles 68% of emails. Automatically.

Marcus, a property manager at Maplewood Properties, spent 90 minutes every morning triaging email. Maintenance requests, lease questions, noise complaints: all arriving in the same inbox, all demanding the same manual read-and-respond cycle. InboxPilot classifies every email the moment it arrives, drafts a reply to each routine request for a person to approve, and surfaces everything else with full context. A manager can switch on automatic sending for verified senders. The inbox that used to be a job became a background process.

COMPOSITE_DEPLOYMENTS Composite scenarios: fictional companies, built on our platform, with outcomes stated as design targets.

BUILT_ON FlowBase · in build

Microsoft Graph FastAPI PostgreSQL Claude Haiku Claude Sonnet Gmail API AWS ECS Fargate
68% emails handled without human input routine maintenance, inquiries, and status updates auto-resolved each week
4 min average response time design target; down from 3.5 hours, priority emails acknowledged in under 2 minutes
11 hrs saved per week across two property managers, time redirected to property inspections and tenant relationships
90 min morning sort eliminated design target; the daily read-and-respond cycle becomes a background process

Property Management Runs on Communication

A property management company lives and dies by how quickly it responds. Not how good the response is, how quickly it arrives.

Property Manager

  • 47 unread emails every morning. Maintenance requests buried between lease renewal questions and parking disputes. No way to know what is urgent until you read every one.
  • 90 minutes of triage before the real work starts, and that is before any inspections, contractor calls, or move-in appointments.
  • Marcus is at a property inspection when Unit 8D emails that their heat is not working. It is Friday at 6:47 PM. He sees it Monday morning.
  • The response took 4 minutes to write. The delay was 63 hours.

Tenant

  • Jamie sent one email: "Heat isn't working in my apartment. It's been like this since this afternoon. Unit 8D." Clear. Specific. Urgent.
  • 63 hours later: "Hi Jamie, thank you for reaching out. We've received your maintenance request and will have someone in touch shortly." The heat was fixed 2 hours after that Monday email was read.
  • The review Jamie left was not about the repair. It was about the silence. "I did not know if anyone received my email. I did not know if it was serious enough to warrant a call. I just did not hear anything."
  • Unacknowledged urgency is the worst tenant experience. Not the problem itself, the void.

Five Stages From Email to Action

Five sequential operations that happen in 4 minutes. Most of the process is invisible to the tenant. They see only the result: acknowledged, helped, resolved.

The InboxPilot Pipeline

1

Receive

Each email reaches FlowBase as a signed webhook, and the pipeline starts before the property manager sees the notification

2

Classify

Claude Haiku produces structured JSON: category, priority, confidence score and its reasoning, all from the email subject and body

3

Judge

A second, independent call to Claude Sonnet reviews the classification against the email and returns a verdict and an adjusted confidence before any action executes. The router acts on that confidence

4

Route

The router compares the checked confidence with the category's thresholds. A high-confidence routine request runs its action at once. Borderline cases go to the review queue with full context. Low confidence, or a matched escalation rule such as a legal keyword, escalates immediately.

5

Execute

The verdict runs the actions the workflow names: a Zendesk or Freshdesk ticket, a Slack message, or an email through Gmail or Outlook to an address the workflow sets. The decision is stored with the judge's verdict and notes. For a routine request the workflow drafts a reply, and it is sent to the tenant who wrote, on the same email thread, once a property manager approves it

Classification Drives the Entire Decision Tree

The classifier prompt doesn't ask for a category. It asks for a structured object: category (maintenance_request | lease_inquiry | noise_complaint | general | legal, this scenario's list), priority (urgent | high | medium | low), confidence (0 to 1.0) and its reasoning. The action is not the model's to choose: the router turns that object into auto-execute, review or escalate from the firm's own thresholds. Every downstream operation (the response tone, the notification trigger, the SLA window) flows from that classification object. If the classification is wrong, the chain produces the wrong result. That's why the judge exists: a second, independent LLM call that checks the classification and adjusts its confidence before the router decides.

The System

Nothing runs on its own until a second model has checked it An email, posted to FlowBase as a signed webhook, is classified by Claude Haiku, and a second model, Claude Sonnet, checks that classification. The router lets an action run on its own only when the checked confidence clears the threshold for that category; everything else waits in the review queue or is escalated. Email as a signedwebhook Classifier Claude Haiku Judge Claude Sonnet Router thresholds percategory Action runs on its own Review queue a person decides Escalation matched by arule holdsescalatestriggersproposes acategorychecksconfidenceacts
Nothing runs on its own until a second model has checked it An email, posted to FlowBase as a signed webhook, is classified by Claude Haiku, and a second model, Claude Sonnet, checks that classification. The router lets an action run on its own only when the checked confidence clears the threshold for that category; everything else waits in the review queue or is escalated. Email as a signedwebhook Classifier Claude Haiku Judge Claude Sonnet Router thresholds percategory Action runs on its own Review queue a person decides Escalation matched by a rule holdsescalatestriggersproposes a categorychecks confidenceacts

The judge checks every classification before the router acts, and the router acts on the judge’s confidence, not the classifier’s.

Four Decisions That Define the Trust

An email automation system that gets things wrong is worse than no automation at all. These four decisions are the difference between a tool that earns production deployment and one that gets turned off after a week.

The judge is not optional

Every decision runs through an independent check before the router acts. The classifier might output 91% confidence. The judge checks whether that confidence is warranted given what the email actually says and the classifier's own reasoning, and the router uses the judge's figure. A misclassified legal complaint that gets a reply nobody approved costs more than a month of operating the entire system. That is the email the judge exists to catch.

Confidence thresholds by category, not globally

Routine maintenance requests get an 80% threshold. Lease inquiries: 88%. Legal or habitability complaints: 100%, so they never run on their own; they escalate to a manager. A single global threshold treats a parking question and an eviction notice identically. They are not identical. The threshold is the risk policy expressed as configuration.

Tone is configuration, not a prompt edit

Each workflow names the tone of its replies in its own configuration, and its escalation rules (legal keywords, a sender writing again) are rows in the database. Warming up Maplewood Properties’ responses after its first month is a configuration change, not a redeploy: no prompt engineering, no risk of breaking the classification chain. Any behaviour that varies by client cannot live inside the model.

A failure lands in front of a person

When a step fails (the model is unreachable, or its answer does not parse), the email is not dropped and nothing is sent. It lands in the review queue as unclassified, with the reason attached, so a property manager sees it the same morning. An automation that fails quietly is how a tenant waits 63 hours again.

The Decision Engine in Practice

Same category. Different confidence. One executes automatically. One waits for human judgement.

Routine Request
Email

"The radiator in my bedroom is making a knocking sound. Can someone take a look?"

category → maintenance_request priority → low confidence → 94% / threshold 80%
JUDGE Kept the confidence at 94%: a routine repair, clearly described.
✓ Auto Clears the 80% threshold: a maintenance ticket opens on its own
Mixed Request
Email

"The radiator you fixed last week is leaking now and there's water on the floor. Is that covered, or should I call a plumber and send you the bill?"

category → maintenance_request priority → high confidence → 73% ← threshold 80%
JUDGE Lowered the confidence to 73%: the email mixes a repair with a question about who pays.
⏸ Review Below the 80% threshold: queued for review

Same category. Different confidence. One executes automatically. One waits for human judgement.

What We Learned

Lessons Learned

The judge is the entire value proposition, not an add-on

Without the judge, you have a classifier acting on its own guess. With the judge, you have a system a property management company trusts with their tenant relationships. One is a script. The other earns a production deployment. We nearly shipped without it to save the cost of that extra model call. That would have been a mistake.

Bad tenant experiences are response time problems, not resolution problems

Jamie's heat was fixed in 2 hours after Marcus read the email Monday morning. The 63-hour wait was the injury, not the repair speed. InboxPilot's core value is closing the acknowledgment gap, not the resolution gap. So this scenario's metrics reward fast acknowledgment, not fast repairs.

Tone is the hardest engineering problem in the system

Every property management company has a different relationship with their tenants. Some are formal; some are warm; some have legal constraints on what they can commit to in writing. Getting tone wrong in a reply is worse than a slow manual response: it creates liability or destroys trust. That is why tone is configuration, set per workflow, rather than written into the prompt.

What I'd Improve

  • A mailbox trigger: FlowBase does not yet watch a mailbox itself, so each email reaches it as a signed webhook today
  • Tenant history weighting: a tenant of 4+ years should carry a higher auto-execute confidence baseline: their communication patterns are known and their requests are rarely edge cases
  • Deep integration with property management software (Buildium, AppFolio) to pull lease status, unit history, and open work orders into the classification context
  • Manager dashboard: category breakdown, response time trends by building, auto-resolve rate, and a weekly summary of review queue patterns to inform threshold adjustments

FlowBase Is the Orchestration Layer

InboxPilot is one configuration of FlowBase, the orchestration layer of the four engines: it takes work in, decides what it is, and sends it where it belongs. The trigger-and-route pattern works the same way whatever starts it: a signed webhook, an authenticated API call, a schedule. A decision below its confidence threshold goes to a person's review queue, and so does a classification that fails, with the reason attached. It also reads DocBase's catalogue of document templates.

Whatever starts the work, FlowBase runs one pipeline A signed webhook, an authenticated API call or a schedule starts the same FlowBase pipeline, which classifies, checks and routes each item to an action or to a person. FlowBase also reads the DocBase template catalogue. Signedwebhook API call authenticated Schedule Pipeline classify, judge,route Action email, Slack,ticket Review queue DocBase templatecatalogue startsstartsstartsactsholdsreadstemplates
Whatever starts the work, FlowBase runs one pipeline A signed webhook, an authenticated API call or a schedule starts the same FlowBase pipeline, which classifies, checks and routes each item to an action or to a person. FlowBase also reads the DocBase template catalogue. Signed webhook API call authenticated Schedule Pipeline classify, judge,route Action email, Slack,ticket Review queue DocBase templatecatalogue startsstartsstartsactsholdsreadstemplates

FlowBase is the orchestration layer. Its one link to another engine today reads DocBase’s templates.

How many hours did your team spend in email this week that didn't need a human?

Book Your Free Assessment