Web Pages is one of the training sources you can give an agent, alongside Video Transcript, Q&A, Conversations (which suggests Q&A rather than training the agent directly), and Documents. Point it at a URL and your agent reads through your site much like a visitor would - opening the page, following the links on it, and pulling in the text it finds - so it can answer questions using your actual documentation, FAQ, and product pages instead of anything you'd have to write out by hand.
You decide how wide that crawl goes and how often it comes back to check for changes, so the agent's knowledge can track a live website without turning into a copy of everything the site contains.
How Web Pages training works
What it is
A Web Pages crawl starts from one URL you give it, discovers other pages linked from that page across the same site, and reads the visible text on each one it's allowed to visit. That text becomes part of what your agent knows.
What it's best for
- Help centers and documentation sites
- FAQ pages
- Pricing and plans pages
- Product or feature pages that explain how something works
How it works
From the Training tab, open Web Pages, enter a starting URL, and click Crawl links. eChat first discovers the pages linked from that URL and shows them to you, so you can choose which paths to keep before anything is trained on. When you're happy with the selection, click Start crawl. The crawl reads the text content of each included page, not the surrounding design.
Advanced options let you set the maximum number of pages to crawl (eChat suggests a limit based on how many pages it found) and turn on Slow scrape for sites that block or throttle fast crawlers.
In Playground and other test modes, a View in Training data link on a Web pages citation opens this source and expands or filters the URL list to the cited page, so you can see exactly which crawl the agent used. See trying it in Playground.
Best practices
Start the crawl from a URL that leads into your content, not your marketing homepage - a help center index or documentation root, for example. That gets the crawl into content-rich pages faster and reduces how much filtering you have to do afterward. If Q&A pairs already cover a topic, remember those take priority over anything a crawled page says, so a Web Pages crawl is best treated as the broad layer underneath your Q&A answers, not a replacement for them.
Scoping the crawl with include and exclude paths
What it is
Rather than crawl every page a site happens to link to, you choose which paths to include and which to exclude. Included paths keep the crawl inside the sections you want; excluded paths keep it out of sections you don't.
How it works
After Crawl links, the discovered paths are split into Included and Excluded lists that you can move paths between. eChat starts you off with a sensible default: your homepage, pricing pages, and language folders are included, help and support paths are included for agents with the Support role, and everything else starts out excluded. Adjust the lists to match your site - for example, include /docs so the crawl covers your documentation, and keep /admin or /login excluded so it never wanders into pages that aren't meant to be read as content. A page in an excluded path is simply never visited.
Best practices
Scope the crawl to the paths that actually hold product and support knowledge - docs, FAQ, and pricing pages are the classic examples. Exclude your blog, press, and marketing pages even if they live on the same domain: they're written to persuade, not to explain precisely, and mixing that tone into training material makes it easier for the agent to answer a support question with something closer to ad copy than fact. Check the default selection before you start the crawl - especially if your agent isn't a Support agent and your help center matters. When in doubt, keep the include list narrow and add more paths later - it's easier to widen a crawl than to untangle answers that came from pages you didn't mean to train on.
Keeping content fresh with re-crawl scheduling
What it is
A re-crawl schedule tells eChat how often to revisit your site and pick up changes, so the agent doesn't keep answering from a snapshot of pages that have since been edited.
How it works
Set Refresh frequency to Weekly (the default), Every 2 weeks, Monthly, or Once (no recrawl). On the schedule you pick, eChat automatically returns to your starting URL and re-crawls the same included paths, re-learning only the pages whose content has changed and picking up new pages that now match your included paths. You can also click Recrawl on the source at any time to refresh it right away.
Best practices
If the source site changes often - pricing updates, documentation edits, new FAQ entries - keep a schedule on so those changes reach your agent without you having to remember to retrain it. A site that rarely changes doesn't need frequent re-crawls, but choosing Once (no recrawl) means you have to click Recrawl yourself whenever the site does change, so it's rarely the better default.
Pages that get skipped
What it is
Not every page the crawler encounters ends up in training. Pages with too little text to be useful, and pages in excluded paths, are left out.
How it works
A page that's mostly navigation, images, or a redirect with barely any body text doesn't add anything an agent could use to answer a question, so it's skipped rather than trained on. Likewise, anything in an excluded path is skipped by design. Skipped pages show a short reason - for example "content too short" or "Excluded by updated path rules" - so you can tell them apart from real failures. Either way, this is expected behavior, not a crawl failure - a lower page count than you expected usually just means your site has fewer text-heavy pages than links.
Crawled pages count toward your plan's training storage. If you reach the limit, the crawl stops adding pages; see training storage.
For eWebinar-linked agents, Web Pages sits alongside a second automatic source, Video Transcript, which syncs your webinar's own transcript without any crawling at all. Between the two, most of an agent's knowledge can come from sources that update themselves once you've set the scope and schedule.