Skip to content
WebChat AI
HomeFeaturesHow it worksIntegrationsPricingDocs
WebChat AI

Build intelligent AI assistants trained on your website content.

Connect with us

Product

  • Features
  • How it works
  • Integrations
  • Pricing
  • Security

Resources

  • Documentation
  • API reference

Legal

  • Privacy Policy
  • Terms of Service

© 2026 WebChat AI. All rights reserved.

Get Started
OverviewQuickstart
Knowledge sourcesFile uploadsRAG & grounding
EmbedCustomizationConfigurationTesting
ConversationsAnalytics & usage
API referenceSecurityTroubleshooting
Changelog

Ready to launch?

Get Started Free
DocsKnowledgeKnowledge Sources

Knowledge Sources & Ingestion Modes

WebChat AI allows you to train assistants from live websites, uploaded documents, or a combination of both.

WebChat AI Source Mode Selector
Select between Website, Documents (Upload-Only), and Website + Docs when creating an assistant.

Source modes comparison

Detailed feature breakdown across the three ingestion modes.

Knowledge source modes comparison
Source ModeWebsite URLFile UploadsBest For
Website
RequiredOptional (via mixed conversion)Public documentation portals, company marketing sites, e-commerce stores, and blogs.
Documents
Not requiredPrimary source (up to 5 files / 10 MB per batch)Internal knowledge bases, customer support runbooks, policy manuals, and technical specifications.
Mixed
RequiredSupported alongside website crawlsCompanies whose public website has base knowledge, but who need to supplement it with offline PDFs or guides.

Website Mode (Automated Crawl)

WebChat AI crawls a public domain or subdomain, following links and indexing pages automatically.

How it works

The crawler discovers pages, strips navigation and boilerplate, and chunks text into 500–800 token passages.

Recommended when: Public documentation portals, company marketing sites, e-commerce stores, and blogs.

Documents Mode (Upload-Only Chatbot)

Create an AI assistant powered strictly by uploaded proprietary files (.pdf, .docx, .md, .txt) without needing a website URL.

How it works

Files are uploaded via multipart form, text is extracted, and passages are vectorized directly into tenant storage.

Recommended when: Internal knowledge bases, customer support runbooks, policy manuals, and technical specifications.

Mixed Mode (Unified Corpus)

Combines content crawled from a public website with additional manual document uploads in a single unified knowledge base.

How it works

Citations link to public URLs for crawled pages and show file download badges for uploaded documents.

Recommended when: Companies whose public website has base knowledge, but who need to supplement it with offline PDFs or guides.

Crawler security & SSRF guard

How WebChat AI ensures secure website ingestion.

When crawling websites, WebChat AI enforces strict network-level security protections:

  • SSRF Protection: Blocks navigation to loopback (127.0.0.1, localhost), private subnet ranges (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16), and cloud metadata endpoints (169.254.169.254).
  • DNS Rebinding Mitigation: Hostnames are re-resolved fresh before every single page fetch and redirect hop.
  • Port Filtering: Non-HTTP ports (SSH, database, SMTP) are rejected.

Related documentation

File Upload Formats & Limits

Learn supported file types, character bounds, and processing statuses.

Read guide

RAG & Retrieval Pipeline

Understand how chunks are embedded, searched, and cited in chat responses.

Read guide

Quickstart Guide

Follow the 7-step guide to connect content and launch the assistant.

Read guide

PreviousQuickstartNext File uploads

Ready to build?

Register a website and get a live assistant in minutes.

Get Started Free