Private beta
— We're onboarding a small group of early-access teams. No credit card required.
Request early access
Capabilities
Every layer of intelligence,
precisely where you need it
From agentic retrieval to SLA enforcement, engineered to handle enterprise complexity without the enterprise headache.
Agentic RAG Engine
Multi-round retrieval that reformulates its own query when the first pass comes up short, fully observable and bounded by configurable rounds.
Phase 5 · Complete
SLA-Aware Routing
Business-hour windows and priority queues route tickets to the right agent at the right moment, with SLA timestamps that stay auditable through policy changes.
Timezone-complete
LLM-as-Judge Evaluation
Every AI response is scored for faithfulness, relevance, and completeness, with hallucination and coverage-gap detection built in, not bolted on.
Automated quality loop
Multi-Tenant Architecture
Complete data isolation per organization, with granular RBAC and invitation flows enforced at the database layer alongside immutable audit trails.
Row-level isolation
Hybrid Retrieval
Vector, keyword, metadata, and graph retrieval fused via Reciprocal Rank Fusion, with semantic deduplication and token-budget compression before the LLM sees anything.
4-stream RRF fusion
BYOK & Provider Flexibility
Bring your own keys for OpenAI, Anthropic, Ollama, or any self-hosted model, with per-org fallback chains and budget hard limits.
Any LLM · Any provider
How it works
From first message to
resolved ticket, automatically
01
Visitor opens the widget
A session is created, the visitor's timezone is detected, and the conversation is routed to the correct widget and team. The QueryPlanner classifies intent and sets retriever activation flags.
02
Parallel retrieval streams launch
Vector, keyword, metadata, and graph retrievers run concurrently against the knowledge base, each contributing candidate chunks scored independently.
03
Results fused and compressed
Reciprocal Rank Fusion merges all streams, near-duplicates are removed, and the surviving chunks are compressed to fit the token budget.
04
Response generated and judged
The LLM produces an answer within the compressed token budget. A separate judge model scores faithfulness, relevance, and completeness, reformulating and retrying if the score falls short.
05
Handoff when it matters
When confidence is low or the visitor requests it, the conversation escalates with full context and a summary pre-loaded for the receiving agent.
Retrieval trace · Live
QueryPlanner.classify()
hybrid + structural
↓
StructuralRetriever
→ boost
↓ parallel launch
Parallel retrieval streams
VectorRetriever
KeywordRetriever
MetadataRetriever
GraphRetriever
↓ ScoreNormalizer → RRF k=60
SemanticDeduplicator
Jaccard 0.85 · 3 removed
ContextCompressor
6 chunks · 4,812 tok
↓ LLM generation
Judge · SUFFICIENT
score 0.91 ✓
Evaluation scores
faithfulness 0.94
relevance 0.89
completeness 0.91
By the numbers
Built for scale from day one
10M+
Messages processed / month
across all tenants
99.9
Uptime SLA, guaranteed
% availability
<2s
End-to-end response latency
p95 across all regions
3×
Faster than traditional helpdesk
avg. projected improvement
Pricing
Transparent pricing,
no surprises
Start free. Scale as your team grows. Every plan includes the full AI engine.
Starter
$15
/ month · $150 / year · 3 users
- 2,000 conversations / month
- 20 documents
- 3 widgets
Most Popular
Pro
$49
/ month · $490 / year · 10 users
- 5,000 conversations / month
- 50 documents
- 10 widgets
Business
$149
/ month · $1490 / year · 50 users
- 10,000 conversations / month
- 200 documents
- 50 widgets
- SSO / SAML
Enterprise
Custom
dedicated infrastructure
- API access
- SSO / SAML
Ready to deploy AI that
actually works?
Private beta · No credit card · Full feature access from day one