Platform

Features

Connect your data sources, index with AI, and search with natural language — self-hosted, permission-aware, and fully customizable.

10
Data connectors
4
LLM providers
3
Reranker options
∞
Future models supported

Data connectors#

Plug-in connector framework — add new sources without touching the core pipeline.

📁

SharePoint

Microsoft Graph API. Sites, document libraries, pages. Full text extraction from PDF, DOCX, PPTX, XLSX, and HTML. Per-document permission crawling.
📝

Confluence

REST API v1. Pages, blog posts, comments, and attachments. Full parent hierarchy traversal and content-level permission reading.
💬

Slack

Coming soon
Bot token integration. Channels, threads, and messages. Historical message indexing with delta sync via timestamp queries.
📂

Google Drive

Coming soon
Service account or OAuth. Documents, spreadsheets, presentations, and shared drives. Native Google Docs content extraction.
📊

Databricks

SQL or Unity Catalog REST API. Index any data stored in Delta tables — call transcripts, CRM exports, ETL pipeline outputs. Configurable column mapping.
📞

Genesys Cloud

OAuth2 Client Credentials. Call transcripts with speaker labels via Analytics API. Region-based URL resolution. Delta sync via timestamp intervals.
🌐

Web Crawler

Index any public website. Respects robots.txt, follows sitemaps, configurable depth and URL patterns.
📦

File Shares (FTP / SMB)

FTP, FTPS, and SMB file servers. Recursive directory traversal with file type filtering and content extraction.
☁️

Cloud Storage

Azure Blob Storage, Amazon S3, and Google Cloud Storage. Unified interface for indexing documents across cloud storage providers.
🗄️

SQL Database

PostgreSQL, MySQL, and SQL Server. Configurable queries to index structured data from relational databases.

RAG pipeline#

Full retrieval-augmented generation pipeline — from raw documents to cited AI answers.

📄

Text extraction

PDF, DOCX, PPTX, XLSX, HTML, and plain text. Handles multi-page documents, tables, and embedded images with OCR support.
✂️

Smart chunking

Three strategies: recursive (general), markdown-aware (wikis), and code-aware (technical docs). Optional contextual chunking prepends document context to each chunk for better embeddings.
🧬

Vector embeddings

Four LLM providers: Azure OpenAI, OpenAI, Anthropic Claude, and Ollama. 1536-dimension vectors stored in pgvector with HNSW indexing. Model fields are free-text — adopt new models without code changes.
🔍

Hybrid retrieval

Vector similarity + keyword search combined. PostgreSQL pgvector for dense retrieval with ACL filtering at query time. Document expansion pulls full context when results cluster from one document.
🎯

Reranking

Three reranker providers: BM25 (default, no external dependency), Cohere Rerank API, and cross-encoder (sentence-transformers).
🔄

Query reformulation

Three strategies to bridge vocabulary gaps: HyDE (hypothetical answer), Multi-Query (alternative phrasings), and Step-Back (broader context). All configurable in Settings — no code changes.
✨

AI generation

LLM generates an answer with inline citations. Server-Sent Events for real-time streaming. Confidence scoring and follow-up suggestions.

Search & Chat#

🔎

Natural language search

Type a question in plain English. Get an AI-generated answer with cited sources, relevance scores, and direct links to source documents. Source filtering by data source, file type, and date range.
💬

Multi-turn chat

Continue any search as a conversation. Full session history with archiving. Each message streams with SSE and includes fresh citations, Knowledge Graph entities, and follow-up suggestions.

Knowledge Graph#

Automatically extract entities and relationships from your documents to surface connections humans would miss.

🕸️

Entity extraction

LLM-powered extraction identifies people, organizations, documents, and custom entity types from indexed content. Runs automatically during crawls with configurable batch sizes.
🔗

Relationship mapping

Entities are linked by relationships extracted from document context — "reports to", "authored by", "references", and more. Weighted edges track relationship strength across multiple sources.
🧩

Community detection

Louvain-based clustering groups related entities into communities with AI-generated summaries. Hierarchical levels let you zoom from broad themes to specific clusters.
🔍

Search & Chat integration

Entities matching your query appear as cards alongside search results and chat answers — showing type, relationships, and document references. Click through to explore the full knowledge graph.
🗺️

Interactive explorer

Browse entities, relationships, and communities from the admin dashboard. Filter by type, search by name, and drill into entity details with all linked documents and relationships.
⚙️

Fully configurable

Enable or disable extraction per data source. Configure entity types, extraction prompts, batch sizes, and community detection thresholds from Settings — no code changes needed.

Search images by meaning, not just filenames — powered by visual AI embeddings.

🖼️

Visual search

CLIP and SigLIP embeddings encode images and text into the same vector space. Search with natural language — "team meeting in a conference room" or "architecture diagram" — and find matching images across all data sources. Smart query preprocessing handles conversational phrasing automatically.
🏷️

AI image analysis

Automatic captioning, classification (photo, chart, screenshot, diagram), semantic tagging, and OCR text extraction. AI-generated metadata enables keyword search alongside visual similarity. Duplicate detection groups near-identical images.

Saved Searches & Alerts#

🔔

Monitor your queries and get notified

Save any search query with a name and frequency (daily, weekly, or after every crawl). AES runs the query on schedule and emails you when new matching content appears. Delta tracking shows exactly what's new since the last check.

How it works
Saved searches use retrieval + reranking only (no LLM generation) for efficiency. The system tracks chunk IDs between runs and highlights new matches with a "NEW" badge. Email alerts include the new content summary and links.

Embeddable widget & API#

🏗️

Embeddable widget

iframe-based widget for embedding search in external web apps. Three modes: floating FAB + slide-out panel, popup window, or new tab. Token exchange auth or mini login form. ~5KB vanilla JS, zero dependencies.
⚡

Answer Snippets API

Programmatic Q&A endpoint for Slack bots, dashboards, and custom integrations. API key auth with per-IP rate limiting. Returns cited answers with confidence scores, token counts, and latency metrics.

Enterprise & administration#

🏠

Self-hosted

Runs entirely in your Docker environment. Your documents, embeddings, and search queries never leave your network.
🔐

SSO & RBAC

Azure AD, Google, and Generic OIDC with IdP group sync. Admin and user roles with permission-aware search results.
🎨

White-label

Your logo, favicon, company name, and accent color across the app, login page, emails, and embedded widget.
🌐

Proxy support

Corporate proxy with CA certificate upload, authentication, and per-target connectivity testing from the diagnostics page.
🩺

Diagnostics

Service health cards, network connectivity checks, app/audit/container log viewer — all in one admin-only dashboard.
📋

Audit trail

Every search query, auth event, settings change, and data source operation recorded, filterable, and exportable.