> For the complete documentation index, see [llms.txt](https://docs.uptiq.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.uptiq.ai/console/knowledge/datasets.md).

# Datasets

Group source files from uploads, cloud drives, or a website crawl.

A dataset groups source files for reuse. It tracks ingestion and accepts new files. Create a knowledge base to process and query the content.

{% hint style="info" %}
Use a dataset when several knowledge bases need the same files. You can also create knowledge directly from attached files.
{% endhint %}

### At a glance

| Task             | Start with                    | Result                            |
| ---------------- | ----------------------------- | --------------------------------- |
| Create a dataset | **Create Dataset**            | A reusable source-file group      |
| Add files        | **Attach Files**              | Files ready to sync               |
| Crawl a site     | **Website** → **Start crawl** | Website content added as one file |
| Replace a file   | **Upload new version**        | The next version of that file     |

{% columns %}
{% column %}

#### Dataset

A named source-file group. It tracks files and their ingestion.
{% endcolumn %}

{% column %}

#### Knowledge base

Processed, queryable content built from datasets or attached files.
{% endcolumn %}
{% endcolumns %}

Inside a knowledge base, **Datasets** shows source files, versions, and sync status. See [Sync & versions](/console/knowledge/sync-and-versions.md).

### Create a dataset

Select **Create Dataset**, then provide the following details:

* **Name** is required. It identifies the dataset.
* **Description** is optional. Use it to distinguish similar datasets.
* **Source Assets** are optional. Add them now or later.

{% tabs %}
{% tab title="Upload" %}
Add files from your computer. Each file can be up to `15 MB`. Add up to `25` files at once.
{% endtab %}

{% tab title="Cloud storage" %}
Add files from **Google Drive**, **OneDrive**, or **SharePoint**. Google Drive uses Google OAuth. OneDrive and SharePoint use Microsoft OAuth.
{% endtab %}

{% tab title="Website" %}
Crawl a public website and add its content to the dataset.
{% endtab %}
{% endtabs %}

{% hint style="info" %}
Datasets support `.pdf`, `.png`, `.jpg`, `.jpeg`, `.webp`, `.tiff`, `.bmp`, and other image formats. Google Drive supports the same formats. Files attached from the agent canvas have a `10 MB` limit.
{% endhint %}

### Add files

To add files later, select **Upload files to dataset** for the dataset. Use **Attach Files** to upload, select a cloud provider, or add a website.

New files show **Ready to sync**. They remain outside knowledge-base answers until you sync them. With automatic sync enabled, syncing begins automatically. See [Sync & versions](/console/knowledge/sync-and-versions.md#sync).

### Add a website

Use **Website** to crawl public pages without downloading them first.

| Setting           | Default | Range     | Effect                                                        |
| ----------------- | ------- | --------- | ------------------------------------------------------------- |
| **Website URL**   | —       | Valid URL | The crawl starting point                                      |
| **Maximum pages** | `25`    | `1`–`200` | Maximum pages ingested                                        |
| **Link depth**    | `2`     | `1`–`10`  | Link levels followed; `1` includes directly linked pages only |

Select **Start crawl**. Crawling runs in the background. The crawled site is added as one file, so it syncs, versions, and detaches as one unit.

{% hint style="warning" %}
Set the page limit before starting. A crawl that reaches its limit contains only part of the site.
{% endhint %}

### Review files and track ingestion

Each dataset lists a file's name, type, size, and ingestion status. Within a knowledge base, files also show their version and sync status. See [File status](/console/knowledge/sync-and-versions.md#file-status).

#### File actions

| Action                     | Result                                                                                                                      |
| -------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| **View document**          | Opens the file                                                                                                              |
| **Upload new version**     | Adds the next file version under the same name. See [File versions](/console/knowledge/sync-and-versions.md#file-versions). |
| **Version history**        | Lists earlier versions                                                                                                      |
| **View extracted content** | Shows processed content. See [Document processing](/console/knowledge/document-processing.md).                              |
| **Download original**      | Downloads the uploaded file                                                                                                 |
| **Detach from knowledge**  | Removes the file from this knowledge base only                                                                              |
| **Delete from dataset**    | Removes the file from the dataset                                                                                           |

| Status        | Meaning                                                     |
| ------------- | ----------------------------------------------------------- |
| **Completed** | Processing finished. The file can contribute after syncing. |
| `in-progress` | Processing is underway.                                     |
| `queued`      | Processing is waiting to start.                             |
| `failed`      | Processing did not finish.                                  |

Wait for **Completed** before syncing. If processing does not complete, wait a few minutes and refresh. Contact your administrator if the issue persists.

### Search and discovery

* Dataset and file search match names and descriptions.
* Search does not inspect document content.
* Use [Try Knowledge](/console/knowledge/try-knowledge.md) to query processed content.

Files are also indexed for semantic discovery. **Preparing for discovery** means indexing is underway. **Discoverable** means indexing is complete. Discovery and sync are independent, so a discoverable file can still be **Ready to sync**.

Discovery searches your data sources by meaning, not by name. Use it to find which dataset holds what you need, for example which one covers this year's rate card:

* On a dataset's page, use **Search this dataset**.
* In the Knowledge Builder chat or the Agent Builder chat, ask for it in your own words.

Only **Discoverable** files are found. A file still **Preparing for discovery** doesn't appear in results yet.

{% hint style="info" %}
**Detaching is not deleting.** Detaching a dataset or file removes it from one knowledge base. Deleting a file removes it from the dataset.
{% endhint %}

### Related

<table data-view="cards"><thead><tr><th>Title</th><th>Description</th><th data-card-target data-type="content-ref">Target</th></tr></thead><tbody><tr><td><strong>Knowledge</strong></td><td>Build queryable knowledge from source files.</td><td><a href="/console/knowledge.md">Knowledge</a></td></tr><tr><td><strong>Sync &#x26; versions</strong></td><td>Bring file updates into a knowledge base.</td><td><a href="/console/knowledge/sync-and-versions.md">Sync &amp; versions</a></td></tr><tr><td><strong>Document processing</strong></td><td>Review processed document content.</td><td><a href="/console/knowledge/document-processing.md">Document processing</a></td></tr><tr><td><strong>Knowledge FAQs</strong></td><td>Find answers about datasets, sync, and versions.</td><td><a href="/console/knowledge/faqs.md">Knowledge: FAQs &amp; Troubleshooting</a></td></tr></tbody></table>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.uptiq.ai/console/knowledge/datasets.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
