step 02

Datasets

Assemble training datasets from the QA pairs auto-generated when your pages were crawled.

01

How it works

When you crawl a page, Kanha automatically generates question-answer pairs from its content, then checks every pair against the page and drops or trims any answer that states something the page does not. These page-level QA pairs are stored on the page and ready to use.

A dataset is an assembly of those existing QA pairs into a single JSONL file for training. Once you create one, Kanha builds it in the background and the dataset page updates itself when it is ready.

Kanha also adds generated examples on top of your pairs: alternative phrasings of your questions, examples that answer from page excerpts the way the live bot does, examples where the excerpts do not contain the answer and the right response is to say so, and a few that teach the bot to introduce itself. They are listed under Generated examples on the dataset page and count toward the pair total.

dataset.jsonl

{"instruction": "What pricing plans do you offer?",
 "response": "We offer three plans: Starter, Pro, and Scale...",
 "system_prompt": "You are a helpful assistant for Example Corp.",
 "source_url": "https://example.com/pricing"}

02

Assemble a dataset

01Choose a site

Go to Dashboard → Datasets and select the site you want to create a dataset from. Make sure the site is verified and has crawled pages with QA pairs.

02Select pages

Optionally select specific pages to include. By default, all indexed pages are included. Use the page selector to exclude pages you don’t want in this dataset.

03Generate

Click Generate. You are taken straight to the dataset page while the build runs in the background. There is nothing to wait around for, the page switches to Completed on its own when the dataset is ready.

Important

Datasets are assembled per site, not per bot. One dataset covers the QA pairs from selected pages on a site. You choose which dataset to use when you start training a bot.

03

Manage QA pairs

Click any completed dataset to see its QA pairs. You can review each question-answer pair and manage which ones are included in training.

01Exclude / include pairs

Toggle individual QA pairs as excluded. Excluded pairs are not used when training or creating new dataset versions. Use this to remove low-quality or irrelevant pairs.

02Filter by page

See which QA pairs came from which page. Click a page in the breakdown to filter the pair list to just that page’s contributions.

04

Dataset versioning

After excluding pairs from a dataset, click Create New Version to snapshot the current non-excluded pairs into a new, immutable dataset. This lets you iterate on your training data while keeping a history of what each bot was trained on.

Each version links back to its parent and shows a version number (v1, v2 and so on). You can train a bot on any version and compare results.

05

Download

Click Download on any completed dataset to get the raw JSONL file. You can inspect it to verify the quality of the QA pairs before starting a training job.

06

Updating datasets

If you update your site content by recrawling pages, adding new pages, or removing old ones, the page-level QA pairs are automatically regenerated on recrawl. Assemble a new dataset to incorporate those changes. Each assembly creates a new dataset entry. You choose which one to use when training.