Platform
Features
Connect your data sources, index with AI, and search with natural language — self-hosted, permission-aware, and fully customizable.
10
Data connectors
4
LLM providers
3
Reranker options
∞
Future models supported
Data connectors#
Plug-in connector framework — add new sources without touching the core pipeline.
📁
SharePoint
Microsoft Graph API. Sites, document libraries, pages. Full text extraction from PDF, DOCX, PPTX, XLSX, and HTML. Per-document permission crawling.
📝
Confluence
REST API v1. Pages, blog posts, comments, and attachments. Full parent hierarchy traversal and content-level permission reading.
💬
Slack
Coming soon
Bot token integration. Channels, threads, and messages. Historical message indexing with delta sync via timestamp queries.
Bot token integration. Channels, threads, and messages. Historical message indexing with delta sync via timestamp queries.
📂
Google Drive
Coming soon
Service account or OAuth. Documents, spreadsheets, presentations, and shared drives. Native Google Docs content extraction.
Service account or OAuth. Documents, spreadsheets, presentations, and shared drives. Native Google Docs content extraction.
📊
Databricks
SQL or Unity Catalog REST API. Index any data stored in Delta tables — call transcripts, CRM exports, ETL pipeline outputs. Configurable column mapping.
📞
Genesys Cloud
OAuth2 Client Credentials. Call transcripts with speaker labels via Analytics API. Region-based URL resolution. Delta sync via timestamp intervals.
🌐
Web Crawler
Index any public website. Respects robots.txt, follows sitemaps, configurable depth and URL patterns.
📦
File Shares (FTP / SMB)
FTP, FTPS, and SMB file servers. Recursive directory traversal with file type filtering and content extraction.
☁️
Cloud Storage
Azure Blob Storage, Amazon S3, and Google Cloud Storage. Unified interface for indexing documents across cloud storage providers.
🗄️
SQL Database
PostgreSQL, MySQL, and SQL Server. Configurable queries to index structured data from relational databases.
RAG pipeline#
Full retrieval-augmented generation pipeline — from raw documents to cited AI answers.
📄
Text extraction
PDF, DOCX, PPTX, XLSX, HTML, and plain text. Handles multi-page documents, tables, and embedded images with OCR support.
✂️
Smart chunking
Three strategies: recursive (general), markdown-aware (wikis), and code-aware (technical docs). Optional contextual chunking prepends document context to each chunk for better embeddings.
🧬
Vector embeddings
Four LLM providers: Azure OpenAI, OpenAI, Anthropic Claude, and Ollama. 1536-dimension vectors stored in pgvector with HNSW indexing. Model fields are free-text — adopt new models without code changes.
🔍
Hybrid retrieval
Vector similarity + keyword search combined. PostgreSQL pgvector for dense retrieval with ACL filtering at query time. Document expansion pulls full context when results cluster from one document.
🎯
Reranking
Three reranker providers: BM25 (default, no external dependency), Cohere Rerank API, and cross-encoder (sentence-transformers).
🔄
Query reformulation
Three strategies to bridge vocabulary gaps: HyDE (hypothetical answer), Multi-Query (alternative phrasings), and Step-Back (broader context). All configurable in Settings — no code changes.
✨
AI generation
LLM generates an answer with inline citations. Server-Sent Events for real-time streaming. Confidence scoring and follow-up suggestions.
Search & Chat#
🔎
Natural language search
Type a question in plain English. Get an AI-generated answer with cited sources, relevance scores, and direct links to source documents. Source filtering by data source, file type, and date range.
💬
Multi-turn chat
Continue any search as a conversation. Full session history with archiving. Each message streams with SSE and includes fresh citations, Knowledge Graph entities, and follow-up suggestions.
Knowledge Graph#
Automatically extract entities and relationships from your documents to surface connections humans would miss.
🕸️
Entity extraction
LLM-powered extraction identifies people, organizations, documents, and custom entity types from indexed content. Runs automatically during crawls with configurable batch sizes.
🔗
Relationship mapping
Entities are linked by relationships extracted from document context — "reports to", "authored by", "references", and more. Weighted edges track relationship strength across multiple sources.
🧩
Community detection
Louvain-based clustering groups related entities into communities with AI-generated summaries. Hierarchical levels let you zoom from broad themes to specific clusters.
🔍
Search & Chat integration
Entities matching your query appear as cards alongside search results and chat answers — showing type, relationships, and document references. Click through to explore the full knowledge graph.
🗺️
Interactive explorer
Browse entities, relationships, and communities from the admin dashboard. Filter by type, search by name, and drill into entity details with all linked documents and relationships.
⚙️
Fully configurable
Enable or disable extraction per data source. Configure entity types, extraction prompts, batch sizes, and community detection thresholds from Settings — no code changes needed.
Image Search#
Search images by meaning, not just filenames — powered by visual AI embeddings.
🖼️
Visual search
CLIP and SigLIP embeddings encode images and text into the same vector space. Search with natural language — "team meeting in a conference room" or "architecture diagram" — and find matching images across all data sources. Smart query preprocessing handles conversational phrasing automatically.
🏷️
AI image analysis
Automatic captioning, classification (photo, chart, screenshot, diagram), semantic tagging, and OCR text extraction. AI-generated metadata enables keyword search alongside visual similarity. Duplicate detection groups near-identical images.
Saved Searches & Alerts#
🔔
Monitor your queries and get notified
Save any search query with a name and frequency (daily, weekly, or after every crawl). AES runs the query on schedule and emails you when new matching content appears. Delta tracking shows exactly what's new since the last check.
How it works
Saved searches use retrieval + reranking only (no LLM generation) for efficiency. The system tracks chunk IDs between runs and highlights new matches with a "NEW" badge. Email alerts include the new content summary and links.
Embeddable widget & API#
🏗️
Embeddable widget
iframe-based widget for embedding search in external web apps. Three modes: floating FAB + slide-out panel, popup window, or new tab. Token exchange auth or mini login form. ~5KB vanilla JS, zero dependencies.
⚡
Answer Snippets API
Programmatic Q&A endpoint for Slack bots, dashboards, and custom integrations. API key auth with per-IP rate limiting. Returns cited answers with confidence scores, token counts, and latency metrics.
Enterprise & administration#
🏠
Self-hosted
Runs entirely in your Docker environment. Your documents, embeddings, and search queries never leave your network.
🔐
SSO & RBAC
Azure AD, Google, and Generic OIDC with IdP group sync. Admin and user roles with permission-aware search results.
🎨
White-label
Your logo, favicon, company name, and accent color across the app, login page, emails, and embedded widget.
🌐
Proxy support
Corporate proxy with CA certificate upload, authentication, and per-target connectivity testing from the diagnostics page.
🩺
Diagnostics
Service health cards, network connectivity checks, app/audit/container log viewer — all in one admin-only dashboard.
📋
Audit trail
Every search query, auth event, settings change, and data source operation recorded, filterable, and exportable.