Help Center/Client guide

Dataset Finder Help Center.

Everything you need to search tens of thousands of curated AI datasets, read dataset cards, organize your workspace, and order custom datasets built and annotated to your specification by our expert labeling team, all from your Dataset Finder account.

Find datasets

Describe what you're building in plain language and get ranked, scored results.

Order custom data & annotation

Don't see what you need? Order a custom dataset — our expert team collects, labels, and annotates it to your spec, delivered with your credits.

Organize work

Save datasets, build projects, and keep your research in one workspace.

01

What is Dataset Finder

Finding the right training data is hard. There are tens of thousands of datasets scattered across Hugging Face, Kaggle, academic pages, and countless other sources, and most of your time goes into sifting through that sea of options, reading papers, and second-guessing quality before you've trained anything at all.

Dataset Finder makes that easy. Instead of hunting through directories and keyword filters, you tell it what you're trying to build, whether that's an AI agent, a fine-tuned LLM, a computer vision model, or anything else, and it surfaces the datasets that actually fit, from a catalogue of tens of thousands of thoroughly curated AI datasets.

But it's more than a search engine. Dataset Finder is your AI training data workspace: a single place to find existing data, understand whether it's any good, and organize it into projects.

And when the data you need doesn't exist yet, you can order it. Through Custom Dataset, you commission bespoke data collection, labeling, and annotation — bounding boxes, segmentation, entity extraction, transcription, preference data, and more — built to your exact specification by Innovatiana's in-house expert annotators. You define the spec and guidelines; we handle the labeling, QA, and delivery. No labeling crowd to manage, no separate vendor, and it all happens inside the same app.

What you can do

Discover

Describe your AI use case in plain language and get curated datasets ranked and scored for relevance, no keyword guessing.

Understand

Every dataset comes with a detailed card: quality scores, licence, biases, limitations, and how to use it, so you don't have to read research papers to judge it.

Organize

Save datasets, group them into AI projects, track experiments and compliance, and start from community AI recipes in your Dataset Workspace.

Commission & annotate

When the data doesn't exist yet, order a custom dataset with your credits. Define the spec and labeling guidelines, and our expert annotators handle collection, labeling, QA, and delivery.

How Dataset Finder handles data

This is worth being explicit about: Dataset Finder does not host, store, or redistribute datasets. We are a curated catalogue and a workspace on top of the data that already exists in the world — not a file repository.

What we maintain is the metadata and the intelligence around each dataset: where it lives, who published it, its licence, size and format, quality scores, biases and limitations, suitable use cases and models. When you bookmark a dataset, your workspace stores that record and your own notes — never the files. ↗ Open Dataset Source sends you to the origin (Hugging Face, an academic page, a government portal, or your own internal storage) to download or access it under its own licence terms.

Curated, not scraped

Datasets are reviewed and enriched by our team rather than dumped in from a crawler, so what you see has been judged worth your time.

Continuously updated

We refresh the catalogue on an ongoing basis — new datasets, corrected metadata, updated licences and scores — so the information doesn't go stale.

You stay in control

No files pass through us, so your own data never leaves your storage. Licence terms remain between you and the dataset's publisher.

You'll see this reflected throughout the app. Adding a dataset by hand says "Only metadata is stored — no file upload needed", and every dataset card carries a Licence field and a Licence check so you can confirm what you're allowed to do with it before you use it. Always review the source's licence yourself before training on it.

The one exception is custom data orders: when you commission a dataset, our team produces and delivers that data to you directly, under whatever terms you agree.

01.1

Create your account

Getting started takes a few steps. You can explore on the Free plan without a credit card.

  1. Sign up with your email address from the Dataset Finder homepage.
  2. Confirm your email. We send you a confirmation link — click it to validate your account.
  3. Complete your profile with your name and a few basic details.
  4. Choose your plan. Start on Free, or pick Startup or Pro if you already know you need more searches and full dataset access.
  5. Add payment details (Startup and Pro only). Billing is handled securely through Stripe. The Free plan never asks for a card.
You can change your plan later from Subscription Plan in the left menu of the app. Upgrading takes effect immediately.
01.2

Choose your plan

Dataset Finder has three plans. The right one depends on how much you search and whether you need full access to every dataset card.

PlanBest forMonthly searchesDataset card access
FreeTrying the product, light explorationGenerous monthly allowanceDescription & Metadata only
StartupSolo builders and small teamsHigher monthly allowanceFull access to every tab
ProActive teams shipping AI productsHigh-volume allowanceFull access + export

For the complete feature-by-feature breakdown, see Compare plans.

01.3

Dashboard tour

When you sign in, you land on your dashboard. The illustration below maps the main areas, numbered to match the list underneath.

Dataset FinderDashboardDataset WorkspaceCustom DatasetSubscription PlanSettingsContact SalesHelp CenterLogoutDescribe what you're building…Ctrl + KPro213 creditsSee our selection of AI datasets — FeaturedDataset 1Dataset 2Dataset 3More SuggestionsTypeDataset NameVolumeFile FormatMatchSaveOpen ↗Dataset titleDescriptionMetadataAnalysisLicenseProviderData TypeCategoryLanguageBusinessEnrichResponsible1234
Dataset Finder dashboard — schematic layout
  • 1 · Search bar — the heart of the app. Describe what you're building and get ranked dataset matches. Press Ctrl + K for the detailed query window.
  • 2 · Left menu — navigate between Dashboard, Dataset Workspace, Custom Dataset, and Subscription Plan.
  • 3 · Credits & plan badge — top right shows your current plan and remaining credits.
  • 4 · Dataset detail panel — selecting a result opens its full dataset card here, with all tabs, Save, and Open Dataset Source.

Two more actions appear on the empty search screen before you run a query: Give it a roll, which surfaces example datasets to explore when you're not sure what to search, and Request a custom dataset, which jumps straight to the custom data request form.

Your plan and credit balance are always visible in the top-right corner so you know where you stand at a glance.
02

How search works

Dataset Finder uses semantic search, not keyword matching. That means you describe your intent — the AI product or model you want to build — and the engine understands meaning, not just the words you typed.

You can search in any style:

  • A vague idea or product concept
  • A precise technical specification
  • An AI use-case description
  • An exact dataset name or keyword (like a known benchmark)

Short or long, technical or plain English, the system reads for intent and returns the datasets that actually fit your need.

Every search returns results — the engine is designed never to come back empty. If nothing matches perfectly, it returns the closest relevant datasets ranked by how well they fit.
02.1

Writing good queries

Because search reads for intent, you can phrase a query however you naturally think about your problem. You don't need to know the exact dataset that exists. Below are the main ways to search, from describing a half-formed idea to naming an exact benchmark.

Describe what you're building

The most powerful queries describe the AI product or model you want to build. Even a vague idea works. If you're not sure what data you need yet, just say what you're trying to make:

"I'm trying to build a medical chatbot"
"I'm building a multilingual customer support copilot for SaaS"
"Train an AI agent to use enterprise SaaS tools end-to-end"
"I want to build a quality-inspection model for a factory line"

Search by use case, model, or training goal

Tell the engine your use case, the kind of model you're working with, or the training method, such as fine-tuning, RLHF, or benchmarking:

"Fine-tune an LLM for legal contract review and clause extraction"
"RLHF datasets for coding models"
"Benchmark dataset for safety testing a chatbot"
"Training data for a recommendation system"

Search by data type or modality

Search across any modality, image, video, audio, text, multimodal, or medical. Mention the type of data you need and what it should contain:

"Audio emotion recognition dataset in multiple languages"
"Satellite image segmentation dataset"
"Multimodal vision-language pairs for document understanding"
"French legal question-answering dataset"

Search by domain, labels, or topic

Name the domain you work in, the labels or taxonomy you need, or a specific topic. The engine understands content, labels, domains, and dataset names:

"Medical imaging dataset with tumor segmentation labels"
"Retail product images with category and brand labels"
"Financial news text labeled for sentiment"

Search by quality, licence, or risk

If usability or compliance matters, say so. You can search by how clean or beginner-friendly a dataset is, its licence, or its privacy risk:

"Beginner-friendly dataset with low privacy risk"
"Commercially licensed dataset, no personal data"
"Clean, ready-to-use dataset that needs minimal preprocessing"

Search by exact name or keyword

Already know the dataset? Search its name or an exact term directly:

"COCO dataset"
"ImageNet object detection"
You can combine all of this in one sentence — domain, modality, use case, model, and any quality or licence needs. For example: "clean, commercially licensed medical QA dataset to fine-tune a healthcare chatbot, low privacy risk." The more intent you give, the more precisely the engine ranks results for you.
02.2

Detailed query mode

For complex needs, open the detailed search request window. Press Ctrl + K (or select the expand icon next to the search bar) to open it. It gives you a larger text area to write a long, natural-language query with multiple criteria at once.

You'll reach a character counter (for example 0/500) showing how much room you have. The limit depends on your plan:

PlanCharacters per query
FreeUp to 140 characters
StartupUp to 500 characters
ProUp to 500 characters

A longer allowance lets you stack context, for example: "Find open-source CCTV datasets from hotels in Europe with at least 1M frames and labeled entrances." Type your request, then select Run search.

Longer queries shine when you have several requirements at once — domain, region, volume, licence, and labeling all in a single description. That's where Startup and Pro's 500-character allowance pays off.
02.3

Reading your results

Every search returns up to 33 results, organized in two groups:

  • Top picks — the 3 strongest matches for your query, shown first as detailed cards.
  • More suggestions — the next 30 results in a scannable list with type, name, volume, file format, and match score.

Each result card shows a small picture, a type icon (image, video, audio, text, medical, or multimodal), a short description, and its relevance score. Select any result to open its full dataset card in the panel on the right.

Want more options after your first search? Use Other suggestions to ask the engine for additional recommendations beyond the original results.
02.4

Relevance scores

Every result carries a relevance score from 0 to 100 in a small bubble. It shows at a glance how well that dataset matches your specific query.

The bubble is colour-coded so you can read it instantly: the closer to 100, the greener it is; the closer to 0, the redder it is.

ScoreColourWhat it means
87GreenStrong match — the dataset closely fits the intent of your query
58AmberModerate match — relevant, but a looser fit worth a closer look
24RedWeak match — adjacent or complementary at best

The colour lets you scan a list of results and spot the strongest candidates without reading every score, then open the green ones first.

Scores are specific to your query. The same dataset can score differently for two different searches, because relevance is measured against what you asked for.
03

The dataset card

Selecting any result opens its dataset card in the right-hand panel. This is the full profile of a dataset, organized into tabs so you can scan exactly what you need.

At the top you'll find the dataset title, its relevance score, a Save button to add it to My Datasets in your workspace, and Open Dataset Source to view or download it from its original location.

The tabs

DescriptionMetadataAnalysisBusiness ApplicationsHow to enrichResponsible AIQ&A
On the Free plan, only Description and Metadata are visible. The remaining tabs show a small lock and blurred preview. See What's locked on Free.
03.1

Description & metadata

The Description and Metadata tabs cover the essentials of what a dataset is and where it comes from.

FieldWhat it tells you
ProviderThe organization or author who built the dataset
LicenseLicense type, such as CC BY, MIT, and so on
Data typeImage, Video, Text, Audio, or Multimodal
CategoryUsage type — Computer Vision, LLM, Image Classification, and more
Main languagePrimary language of the data (en, fr, es, …)
DomainField such as medical, legal, finance, geospatial, retail
File formatJSONL, CSV, Parquet, PNG, and others — multiple are possible
Data sizeNumber of rows, files, or samples (e.g. 35,000 images)
Storage sizeApproximate file size in MB

You may also see up to three thumbnail images illustrating the dataset, plus the source URL used by Open Dataset Source.

03.2

Analysis & scores

The Analysis tab is where Dataset Finder goes beyond a basic listing. Each dataset is rated on quality dimensions with a star rating out of 5, each with an explanatory comment, so you can judge usability before you commit.

Ease of use

Hard to work with … Very easy

How hard the dataset is to work with, from 1 star (hard to work with) to 5 stars (very easy).

Cleaning needed

Not clean at all … Super clean

How much data cleaning is required, from 1 star (not clean at all) to 5 stars (super clean, no cleaning needed).

Label richness

No labels … Very precise

The quality and depth of labels, from 1 star (no labels) to 5 stars (very precise and detailed labels).

Other analysis fields

  • List of labels — the labels available, with counts and descriptions
  • Class distribution — the share of each class, when available (e.g. 90% cats, 10% dogs)
  • Beginner friendly — whether it suits less experienced users
  • Source reliability — verified, user-uploaded, academic, government, or private
  • PII content — whether personal data is present, confirmed or potential
  • Dimensionality — the shape of the data (e.g. 50×50 images, paragraph-level text)
  • Fine-tuning ready — guidance on which training it suits
03.3

Business applications

The Business Applications tab helps you see how a dataset translates into real AI work.

  • Primary use cases — the AI use cases where this dataset is useful
  • Useful for models — existing models you can train or fine-tune with it (for example YOLO, or fine-tuning Mistral)
  • Known limitations — constraints to keep in mind when applying it

How to enrich

A separate How to enrich tab gives recommendations for improving the dataset: cleanup, balancing, adding labels, or complementary datasets to combine it with. If a dataset is close but not quite enough on its own, this is where you'll find ideas for closing the gap, including when a custom dataset could fill what's missing.

03.4

Responsible AI

Responsible AI is about building models that are fair, safe, transparent, and compliant. A model is only as good as the data it learns from, so a lot of responsible-AI work starts at the dataset: if the data is biased, unrepresentative, or carries legal or privacy problems, those issues get baked into your model and surface later as unfair outcomes, compliance gaps, or reputational risk.

The Responsible AI tab surfaces those considerations up front, so you can make an informed choice before a dataset ever reaches training. Instead of discovering a problem after deployment, you see the known risks while you're still evaluating candidates.

What the tab shows

  • Known biases — biases documented or discussed for this dataset, such as skew toward certain groups, classes, or conditions that could make a model treat some inputs unfairly
  • Cultural diversity — representation gaps, for example data drawn from a single region, language, or demographic, which limits how well a model generalizes
  • Known limitations — documented constraints the community has raised, such as label noise, collection artifacts, or narrow coverage
  • Ethical concerns — concerns raised online or in published discussion, including how the data was collected and whether consent or sourcing is in question

Why it matters

Regulations like the EU AI Act increasingly expect teams to understand and document the data behind their models, especially for higher-risk use cases. Reviewing this tab helps you:

  • Catch bias early, before it becomes unfair model behaviour in production
  • Judge fit for your specific users and regions
  • Document your due diligence for compliance and governance
  • Decide whether to enrich or rebalance the data (see the How to enrich tab)
Always review this tab before training on a dataset for a production or regulated use case. Bias, PII, and licence terms can affect whether a dataset is safe and compliant for your application. These signals are drawn from public documentation and community discussion, so treat them as a starting point for your own review, not a guarantee.
03.5

Q&A

The Q&A tab answers common questions about a dataset in a quick, scannable format: pairs of questions and answers covering the practical things people usually want to know before using it.

It's the fastest way to get specific answers without reading every other tab. Typical questions cover things like what the dataset contains, how it's licensed, whether it needs cleaning, what it's best suited for, and any catches to be aware of.

When you save a dataset to My Datasets, you can also add your own Q&A entries to it, each a question and its answer, to capture notes for yourself or your team. Use + Add Q&A on the dataset to add an entry.

Use Q&A as a final sanity check once a dataset looks promising from its Description, Analysis, and Responsible AI tabs, then add your own questions as they come up while you work with it.
04

Dataset Workspace overview

The Dataset Workspace is where you and your team manage everything you're working with. Open it from Workspace in the left menu.

The pill at the top left is your workspace switcher. Open it to see every workspace you belong to, switch between them, Manage members, or Create workspace. You can run several — one per team, client, or product line — and each one carries its own datasets, AI projects and recipes.

Next to the switcher are the two workspace tabs, My Datasets and My AI Projects, each showing a live count. AI Recipes, the shared community library, opens from the button on the right.

INNOVATIANA AI TEAM ▾My Datasets 6My AI Projects 18AI RecipesSearch my datasets…Status ▾Source ▾Data Type ▾6 datasets+ Add DatasetMY DATASETS (6)Customer Support Chat Logs2.8M conversationsCANDIDATEPrivateRetail Review SentimentParquetIN USEDatasetFinderWildlife Image Classification25,000 imagesCANDIDATE1 projectCustomer Support Chat Logs v20262.8 Million conversations (~185 GB)Link to Project +↗ Open SourceCANDIDATE ▾Private / InternalTextDescriptionMetadataAnalysisResponsible AIQ&AMy NotesProjects✎ EditPROVIDERDATA TYPETextLICENSESOURCE URLLICENSE CHECKNot assessed
The Dataset Workspace: workspace switcher, My Datasets and My AI Projects tabs, the AI Recipes library, and a dataset's detail panel

My Datasets

Every dataset your workspace has bookmarked from search or added by hand, each with its own status, notes, and tags so the team can see where it stands.

My AI Projects

Shared project records: model, linked datasets, evaluation, compliance, documentation, and a project health score.

AI Recipes

A community library of end-to-end blueprints you can browse, vote on, save, clone into your workspace, and publish back.

The three work together: you bookmark datasets into My Datasets, build them into My AI Projects with evaluation and compliance tracking, and start new work by cloning a blueprint from AI Recipes.

Dataset Finder does not host or store datasets. Your workspace holds metadata only — names, sources, licences, analysis, your statuses and notes. There is no file upload, and the data itself always stays at its origin (Hugging Face, your own storage, or wherever it lives). ↗ Open Dataset Source takes you to it. See How Dataset Finder handles data.

Workspaces & your team

A workspace is a shared container for your team's research. Rather than everyone keeping a private list of bookmarks, you create a workspace and invite the people you work with into it.

  • Create workspace — from the workspace switcher, set one up for a team, a client, or a product line. You can belong to several and switch between them at any time.
  • Invite team members — open Manage members from the same menu to invite people and see who's in each workspace. The switcher shows the member count next to every workspace.
  • Roles — each workspace shows your role on it. An Owner manages the workspace and its members; a Viewer can see the shared datasets, projects, and recipes.
  • Work from a shared library — bookmarks, statuses, notes, and tags are visible to everyone in the workspace, so two people don't evaluate the same dataset twice.
  • Keep projects together — AI projects belong to the workspace, not to one person, so work continues when someone is away or changes role.

What your team can track

Bookmarked datasets

Every dataset saved from search lands in the workspace library, with status (Candidate, In use, and so on), notes, and tags your team maintains together.

Cloned recipes

Use ⧉ Use this Recipe to clone a community blueprint into your workspace and adapt it — datasets, model config, and workflow come with it.

All plans can bookmark and manage datasets in My Datasets. My AI Projects and team seats scale with your plan: Free includes one demo project, Startup up to five, and Pro unlimited or high-cap projects with full documentation and export. See Compare plans.
04.1

My Datasets

The My Datasets tab is your workspace's inventory of datasets worth tracking, shared with everyone you've invited. Datasets arrive here two ways:

  • Bookmark from search — select Save on any dataset card and it appears here for the whole workspace.
  • Add manually — select + Add Dataset to enter one yourself.

Adding a dataset

+ Add Dataset opens a short form. It states it plainly at the top: only metadata is stored — no file upload needed. You're registering a pointer to the data plus your own notes, not moving the data anywhere.

Choose a source type and say where the data lives:

  • Hugging Face — paste the dataset path (for example teknium/OpenHermes-2.5 or datasets/squad)
  • Private / Internal — give an internal path or endpoint for data you host yourself

Then fill in a dataset name (required), and optionally its Status, Data Type, a short description, and any notes on preprocessing, limitations, or known issues. Select Add to Workspace to save it.

Afterwards you can enrich the record with the same catalogue fields you see on a dataset card: provider, licence, category, language, domain relevance, keywords, file format, sizes, dimensionality, list of labels, class distribution, source reliability, PII content, quality scores, use cases, models, limitations, biases, and a Q&A. Anything you don't know can be left blank.

Your own fields

On top of the catalogue data, every saved dataset carries fields that are yours to manage:

FieldWhat it's for
StatusWhere the dataset stands in your work (e.g. in use, candidate, approved, flagged)
NotesYour own notes on the dataset
TagsLabels to organize and filter

Finding datasets in your library

As your library grows, narrow it down with the bar across the top: Search my datasets… for a free-text match, then filter by Status, Source, and Data Type. The dataset count sits on the right, next to + Add Dataset.

Working with a saved dataset

Open any dataset to see its full record on the right. It's organized into tabs — Description, Metadata, Analysis, Responsible AI, Q&A, My Notes, and Projects — with a badge row showing its status (a dropdown you can change), its source (Private / Internal or DatasetFinder), and its data type.

From the header you can ✎ Edit its fields, + Link to Project to attach it to one of your AI projects, ↗ Open Dataset Source to go to the data at its origin, or delete it from the workspace.

You edit your own fields and any details you entered, but datasets pulled from the catalogue keep their published information intact.
Mark a dataset as flagged if you spot a problem. Flagged datasets are surfaced in your project health checks so issues don't slip through.
04.2

My AI Projects

The My AI Projects tab turns datasets into managed projects. A project is one AI system, with everything it needs in one record: the model, the datasets it uses, experiments, evaluation, compliance, and documentation. Select + New Project to create one.

Setting up a project

When you create a project you capture the essentials:

  • Project name and the AI model — base, fine-tuned, or custom, with its model name / identifier
  • Risk level — how sensitive the system is
  • Lifecycle stage — for example development, staging, or production
  • Deployment, owner, and an experiment tracking link
  • Eval metrics & thresholds and a retraining trigger

What a project record contains

Each project is organized into sections you can fill in over time:

SectionWhat it holds
Project OverviewModel, risk level, lifecycle, owner, deployment, and the project's purpose
Evaluation & PerformanceExperiments, metrics, thresholds, and how the model is performing (see Experiments & tracking)
Known Failure ModesDocumented ways the model can fail, with severity and mitigations
Compliance & GovernanceEU AI Act and governance fields (see Health & compliance)
Training Data InventoryThe datasets linked to this project, each with a usage role
Documentation NotesDecision records, incidents, changelog, and external references

Linking datasets

Use + Link Dataset to attach a dataset from My Datasets to the project, then assign its usage role and optional usage notes:

TrainingValidationTesting

The linked datasets make up the project's Training Data Inventory, so you always know exactly what data went into the model and in what role.

Documenting your work

The Documentation Notes section keeps an audit trail of how the project evolved:

  • Architecture Decision Records (ADRs) — log a decision with its title, rationale, and date
  • Known Failure Modes — record a failure with severity, description, and resolution / mitigation
  • Incident Log — log incidents as they happen
  • Changelog — version entries with author / approver and a summary of changes
  • External References — links out to related resources, with a label and URL
The number of projects you can create depends on your plan: one demo project on Free, up to five on Startup, and unlimited or high-cap on Pro. Full documentation export is a Pro feature.
04.3

Health & compliance

Each project shows two scores so you always know its state at a glance.

Health score

A 0 to 100 score with a status label, computed from your project's own data. It checks things like:

  • Do all evaluation metrics have a passing result against their thresholds?
  • Are the compliance fields filled in?
  • Are there any open incidents?
  • Has the project been evaluated recently (within 60 days if it's in production or staging)?
  • Are any linked datasets flagged?

For example, a healthy project might show 92 while one with open issues shows 61.

Compliance score & the EU AI Act

The compliance score is a percentage based on how many key governance fields you've completed. The project's Compliance & Governance section is built around the EU AI Act and includes fields such as:

  • Risk level and regulatory category (EU AI Act Annex III)
  • Intended users, affected persons, and geographic scope
  • Human oversight mechanism (Art. 14)
  • Post-market monitoring plan (Art. 72)
  • Eval metrics & thresholds, last evaluated date, and retraining trigger criteria
  • Inference environment and deployment scope

The score comes with a list of exactly what's still missing, so you know what to complete next.

AI Compliance Readiness Report

Select ⚖ AI Compliance Report to generate a consolidated AI Compliance Readiness Report. You choose which sections to include, then export it as PDF or CSV, useful for demonstrating data governance and due diligence to stakeholders, auditors, or customers.

The report pulls together overview, evaluation, failure modes, compliance, training-data inventory, and documentation into a single document you can hand to a reviewer.
04.4

Experiments & tracking

Within a project's Evaluation & Performance section, experiments track your model runs and how they perform over time.

Each experiment

  • Has an experiment name, the model used, a last evaluated date, and a set of metrics
  • Records metrics & thresholds — each metric can be checked against a target to show whether it's passing
  • Can be marked with ★ Set as main model. The main experiment's model and metrics show on the project, so it always reflects your current state

Snapshots and trends

Select 📸 Snapshot to save the current metrics as a point in history. Snapshots build up over time into a trend, so you can see whether the model is improving run over run. Use + Add metric to track a new measure and 💾 Save to store your changes.

Metrics with thresholds feed directly into the project health score — a project with failing or stale metrics scores lower, prompting a re-evaluation.
04.5

AI Recipes

The AI Recipes tab is a community library of end-to-end blueprints for building AI use cases. A recipe describes the problem, the data, the model, the workflow, and how to evaluate it, so you can start from a proven approach instead of a blank page.

What a recipe contains

  • Overview — title, short description, who created it, industry, primary use case, data modality, difficulty, and time to first prototype
  • Problem & output — the business or research problem and the expected output
  • Model & approach — model family, specific model / checkpoint, training strategy, deployment target, recommended stack and tooling
  • Workflow steps — the steps to follow, in order
  • Dataset ingredients — the datasets the recipe calls for
  • Evaluation metrics — how to measure success
  • Constraints — risk level, latency requirement, and whether it's a regulated environment

Browsing and filtering

Search with Search recipes…, then filter the library by Data Type, Use Case, Industry, Model Family, Training Strategy, and Difficulty. ♡ Saved narrows to recipes you've saved, and ⇅ Sort changes the order (for example Top). The recipe count sits on the right.

Each card shows its difficulty, use-case and industry tags, the author, time to prototype, and its vote, view and comment counts. Open one to see the full recipe in Overview, Datasets, Model & Training, Evaluation, How-to, Products, and Comments tabs.

Using a recipe

  • ⧉ Use this Recipeclone the blueprint into your workspace, pre-filled and ready to adapt. The datasets, model config, and workflow steps come with it, and your copy is shared with everyone in the workspace.
  • Vote — upvote or downvote recipes to help the community surface the best ones
  • ♡ Save — keep a recipe for later; the ♡ Saved filter shows just those
  • Comment — discuss a recipe, reply to others, and flag content; you can edit or delete your own comments

Publishing your own

Share your approach by publishing a recipe. Fill in the recipe fields, add workflow steps, dataset ingredients, and metrics, then publish it to the community. You can Unpublish it later if you want to take it down.

Using a recipe is the fastest way to begin: it sets up the model, datasets, and workflow from a blueprint that has already worked for someone else.
05

Order a custom dataset & annotation

When the data you need doesn't exist yet, order it. The Custom Dataset section is the data labeling and annotation side of Dataset Finder: instead of searching for an existing dataset, you have one built to your exact specification by expert human annotators.

This covers the full range of annotation work — bounding boxes, polygons, segmentation masks, keypoints, entity spans, classification, transcription, and ranking or preference data — across computer vision, NLP, document processing, Gen-AI, content moderation, and medical data. You can also request a dedicated annotation platform for your project if you need one.

Open Custom Dataset from the left menu, then select New Request to start. Your existing orders live under My Datasets Orders in the same place.

Two ways it works

A custom request can go one of two ways:

We source it for you

In some cases we can find or assemble the data you need on your behalf, then label it to your spec.

You bring the raw data

Most often, you upload your raw data plus your annotation guidelines, and our team labels it exactly as you instruct.

What happens after you submit

  1. Credits are deducted when you submit, based on the complexity and volume of what you ordered. You're told if you don't have enough credits before you submit.
  2. You get a delivery date. Even the largest requests are typically delivered within a month.
  3. For very large or ambitious requests, we'll reach out to set up an Enterprise contract instead, so we can optimise the effort, your costs, and the delivery timeline together.
The number of credits a request consumes depends on its complexity and volume. You'll always see whether you have enough before submitting, and you can buy more credits if needed.
05.1

Filling the request form

The request form captures everything our team needs to build your dataset. Required fields are marked with an asterisk.

New RequestMy Datasets OrdersBuy more creditsSubmitData type *Select data typeData volume *Enter volumeUse case *Select use caseSearch…ImageVideoAudioTextMultimodalMedical DataTaxonomy *Number of classesEnter numberNumber of objects per itemEnter numberI need an annotationplatform set up…Attach guidelines *Upload your brief or dataset specs (PDF, Zip or Word, <= 20 MB)Upload filesI've already uploaded my dataset to my own platformUpload Raw Dataset *Upload your dataset here (max 5 GB; text, image, audio or video)Upload1234567
New Request form — schematic layout
#FieldDetails
1Data type *Image, Video, Audio, Text, Multimodal, or Medical data. The dropdown is searchable.
2Data volume *How much data you need (numerical, up to 5 digits)
3Use case *Computer Vision · NLP / Text Classification · NLP / NER (Entity Extraction) · Document Processing (OCR, key-value extraction) · Gen-AI dataset (prompt-output pairs, instruction tuning) · Content Moderation · Reinforcement Learning (preference data, trajectories, AI feedback) · Medical Imaging / Clinical Data · Medical Text Classification / Annotation (notes, reports, EMR)
4Taxonomy *Number of classes and number of objects per item
5Annotation platformToggle on if you need an annotation platform set up for your project
6 · 7Guidelines & raw data *Upload areas for your brief and your raw dataset — covered in Guidelines & uploads
Turning on "I need an annotation platform set up for my project" notifies our team and adds a credit overhead, because it includes standing up the labeling environment for you.

Choosing your use case

The Use case dropdown tells our team what kind of model you're building, which shapes how the data is structured and labeled. Here's what each option means and the kind of training data it typically needs:

Use caseWhat it isTraining data typically needed
Computer VisionModels that interpret images or video, such as object detection, image classification, segmentation, or tracking.Images or video frames with labels: bounding boxes, polygons, segmentation masks, keypoints, or class tags per item.
NLP / Text ClassificationModels that sort text into categories, such as sentiment analysis, topic tagging, intent detection, or spam filtering.Text samples (sentences, reviews, tickets, posts) each labeled with one or more category tags from your taxonomy.
NLP / NER (Entity Extraction)Models that find and tag specific entities inside text, such as names, dates, amounts, products, or domain-specific terms.Text with span-level annotations: each entity marked by its start and end position and an entity type label.
Document Processing (OCR, key-value extraction)Models that read structured or semi-structured documents, such as invoices, receipts, forms, and contracts, and pull out fields.Document images or PDFs with transcribed text and labeled key-value pairs or regions (for example invoice number, total, date).
Gen-AI dataset (prompt-output pairs, instruction tuning)Data to fine-tune or instruction-tune generative models, so they follow instructions and respond in the style and format you want.Prompt-and-response pairs, instruction-input-output triples, or multi-turn conversations written or curated to your guidelines.
Content ModerationModels that flag unsafe, harmful, or policy-violating content across text, image, audio, or video.Examples labeled against your moderation policy: safe vs. violating, with category tags (e.g. hate, violence, spam) and severity.
Reinforcement Learning (preference data, trajectories, AI feedback)Data to align models with human preferences (RLHF) or train agents, by comparing outputs or recording action sequences.Ranked or pairwise preference comparisons between model outputs, reward labels, or recorded trajectories of states and actions.
Medical Imaging / Clinical Data (vision-oriented)Vision models for healthcare, such as detecting findings in X-rays, CT, MRI, pathology slides, or other clinical images.Medical images annotated by qualified annotators: segmentation, bounding boxes, or classifications, often with clinical context.
Medical Text Classification / Annotation (notes, reports, EMR)Models that read clinical text, such as notes, radiology reports, and electronic medical records, to classify or extract information.Clinical text labeled with categories or entity spans (diagnoses, medications, procedures), handled with appropriate privacy care.
Not sure which fits? Pick the closest match and describe the specifics in your guidelines. If your project spans more than one (say, document processing plus entity extraction), note that too, our team will scope it with you.
05.2

Guidelines & uploads

Two uploads help our team build exactly what you want: your guidelines and, if you have it, your raw data.

Attach guidelines

Upload your brief or dataset specs as a PDF, Zip, or Word file (up to 20 MB). Your guidelines are the single most important factor in getting a dataset you can actually use. They tell our labelers exactly what "correct" looks like, so the more precise they are, the better the result and the fewer revision cycles you'll need.

How to write great labeling guidelines

Good guidelines remove ambiguity. A labeler should be able to read them and label an item the same way you would, without guessing. Here's what to include:

1. State the goal and the use case

Open with one or two sentences on what the dataset is for. Labelers make better judgement calls when they understand the intent.

Example: "This data will train a model to detect defective products on a factory line. When unsure, prefer flagging a possible defect over missing one."

2. Define every label precisely

For each label or class, give a clear definition in plain language. Don't assume the name is self-explanatory. Spell out exactly what does and does not belong in each category.

  • Label name — exactly as it should appear
  • Definition — what it means, in one or two sentences
  • Include — what counts as this label
  • Exclude — what looks similar but should not get this label

3. Show examples, including the hard ones

Examples are worth more than definitions. For each label, give a few positive examples, and crucially a few edge cases and near-misses. The items labelers disagree on are almost always the ambiguous ones, so decide those for them in advance.

  • Clear positives — obvious cases of each label
  • Edge cases — borderline items and how you want them handled
  • Negatives / near-misses — things that look right but aren't
  • Counter-examples — common mistakes to avoid

For image, video, or document work, an annotated example next to a wrong one resolves more confusion than any description. Take a bounding-box task: the difference between "correct" and "wrong" is obvious in a picture, but hard to pin down in words.

CorrectToo looseCuts off objectMissed entirely
One correct annotation vs. three common mistakes

A single labelled diagram like this, showing a tight, correct box beside boxes that are too loose, that cut off the object, or that miss it entirely, tells a labeler exactly what you expect. Do the same for whichever task you have: spans for text, masks for segmentation, fields for documents.

Sample annotated data is the single most valuable thing you can provide. Even a handful of correctly labelled examples, often called a gold set, does more to align our team with your intent than pages of written rules. If you can label just 10 to 20 items yourself and include them with your guidelines, your dataset will come back far closer to what you wanted, with fewer revision cycles.

4. Give rules for ambiguity and conflicts

Tell labelers what to do when an item is unclear or could fit more than one label. A few decisions to make for them:

  • What to do when two labels could both apply (pick one? allow multiple? a priority order?)
  • What to do with low-quality, blurry, or partial items
  • Whether there's a "none of the above" / "unsure" option, and when to use it
  • How to handle items in a different language or outside the expected scope

5. Specify the format and taxonomy

Be explicit about the shape of the output so it matches what your pipeline expects:

  • Label set — the full list of allowed labels or classes
  • Annotation type — e.g. bounding boxes, polygons, segmentation masks, keypoints, spans, classification, transcription, ranking
  • Granularity — per item, per region, per token, per frame
  • Multiple labels — whether an item can have more than one
  • Output format — JSONL, CSV, COCO, and any field names you need

6. Note quality and compliance requirements

If you have specific quality bars or sensitivities, say so up front:

  • Target accuracy or agreement level, if you have one
  • Any PII or sensitive content to redact, skip, or handle carefully
  • Domain expertise required (for example medical, legal, or a specific language)
  • Anything that is a hard "do not" for your project
A simple structure that works: 1) goal and use case · 2) the label set · 3) one definition + examples block per label · 4) edge-case and conflict rules · 5) output format · 6) quality and compliance notes. Start small and add edge cases as you think of them.

Upload raw dataset

If you already have raw data to be labeled, drag and drop it here (up to 5 GB; text, image, audio, or video). If you don't have data yet, that's fine, our team can source it as part of your request.

If your dataset is too large to upload, or hosted elsewhere, tick "I've already uploaded my dataset to my own platform" and share the access link in your guidelines. Don't forget to actually grant access, or work can't start.

Use your own labeling platform

Already have a labeling environment set up with your own vendor? We can work directly inside it. If your team uses a tool like Encord, Label Studio, CVAT, Labelbox, V7, SuperAnnotate, Roboflow, or any similar platform, our annotators can plug straight into your existing setup instead of using ours.

To integrate, you just share access:

  • Invite our team to your project, or provide accounts / seats on your platform
  • Point us to the data and the label schema already configured there
  • Note any platform-specific instructions in your guidelines (project names, workflows, review steps)

Mention this in your guidelines, or turn on "I need an annotation platform set up for my project" in the request form if you'd rather we set one up for you instead.

Working in your own platform keeps the labeled data, audit trail, and review workflow entirely within your environment, which is helpful when you have existing pipelines or specific governance requirements.

When everything's ready and you have enough credits, select Submit. You'll get a confirmation that your request has been received.

05.3

Tracking your orders

Open My Datasets Orders to see all your requests. Each row is one order, with its order number, data volume, data type, the guidelines you attached, its current status, and the actions available.

New RequestMy Datasets OrdersOrder NoData VolumeData TypeGuidelinesStatusActionCompletedDownloadCompletedDownloadRequested×CompletedDownloadQA ReviewRejected1
My Datasets Orders — schematic layout

Order statuses

The Status column tells you where each order stands:

StatusWhat it means
RequestedSubmitted and waiting to be validated by our team. This is the only stage where you can still cancel.
QA ReviewLabeling is done and the dataset is going through quality checks.
CompletedReady. The Download button appears so you can collect your dataset.
RejectedThe request couldn't be fulfilled as submitted. Open the details to see why, or contact our team.

Actions on each order

The Action column always includes a small eye icon for viewing details. Download and Cancel appear depending on the order's status.

  • View (the eye icon) — available on every order. Opens an Order Details panel with the full recap: data type, volume, use case, taxonomy (number of classes and objects per item), annotation-platform choice, submission date, and the guidelines file.
  • Download — appears once an order is Completed, to collect the finished dataset.
  • Cancel (the red ×) — only shown while an order is still Requested.
Hover an in-progress order's status to see its estimated delivery date. Even the largest requests are typically delivered within a month.
You can cancel an order only while it's still Requested. Once our team has validated it and started work, cancellation is no longer available.
06

How credits work

Credits are the currency for custom dataset orders. The flow is simple: you hold a credit balance, you spend credits when you order a custom dataset, and you top up whenever you need more.

  1. Get credits — buy a credit package, or use the monthly credits included with your plan.
  2. Order a custom dataset — when you submit a request, credits are deducted based on its complexity and volume.
  3. Top up anytime — if a request needs more than you hold, you're told before submitting, and you can buy more.
  • The exact credit cost is calculated automatically from each request.
  • Your current balance is always visible in the top-right corner of the app.
The Pro plan includes 20 credits per month as part of the subscription. These refresh monthly and can be used toward custom dataset orders, on top of any credit packages you buy.
06.1

Buying credits

Select Get credits (or Buy more credits) from the dashboard, the Subscription Plan page, or inside the custom request form. You'll see a set of credit packages plus a custom option.

100 creditsEUR 75Get credits500 creditsEUR 375Get credits1000 creditsEUR 750MOST POPULARGet creditsCustom creditsEUR 1,000+Get credits-2500+iInformation123
Get credits — schematic layout
  • 1 · Preset packages — fixed bundles of credits at a set price, from small (for POCs and experiments) up to large (for production-scale pipelines). One is marked Most popular.
  • 2 · Custom credits — for large-scale projects, set your own amount with the stepper. The estimated price and how many labeled items it covers update as you adjust it.
  • 3 · Information — a note on what credits unlock and how to reach the team for a tailored quotation.

Pick a package and select Get credits to add them to your balance. The more credits in a package, the more labeled items they cover.

Plan discounts on custom orders

Your subscription tier affects how far your credits go when ordering custom datasets:

PlanDiscount on custom dataset credits
FreeNo discount
Startup10% discount
Pro25% discount, plus 20 included credits each month
All plans, including Free, can buy credits and place custom dataset orders. Higher tiers simply get a better rate, and Pro adds monthly included credits. For very large projects, the team can prepare a tailored quotation, see Request a custom dataset.
06.2

Subscriptions & billing

Subscriptions are billed securely through Stripe. Manage your plan and payment details from Subscription Plan in the left menu, or by selecting the credit-card icon in the top bar.

  • Free requires no card and no payment details.
  • Startup and Pro require a card on file via Stripe.
  • You can switch plans at any time; changes take effect immediately.
07

Compare plans

The full feature comparison across all three plans:

FeatureFreeStartupPro
Dataset searchYesYesYes
Semantic searchFullFullFull
Characters per query140500500
Dataset card accessPartialFullFull
Advanced metadataLocked / previewFullFull
Save datasetsUp to 10YesYes
WorkspaceBasicAdvancedAdvanced
Number of projects1 demo5Unlimited / high cap
Dataset notesLimitedYesYes
Export / documentationNoNoFull export
Custom dataset requestsYesYesYes
Custom dataset discountNo10%25%
Buy creditsYesYesYes
Card requiredNoYesYes
07.1

Search limits

Each plan includes a monthly search allowance plus a daily anti-abuse cap. If you hit the daily cap, there's a short cooldown before you can search again.

PlanDaily capCooldown
Free10 / day2 hours
Startup100 / day2 hours
Pro250 / day2 hours

If you reach a limit, you'll see a clear message. For example, on the daily cap:

"You've reached today's search limit. You can search again in 2 hours, or upgrade to continue researching now."

And when your monthly allowance is used up, you'll see when it resets and how to continue right away by upgrading.

Your monthly allowance resets at the start of each month. If you regularly hit your limit, upgrading to the next tier gives you a higher allowance.
07.2

What's locked on Free

The Free plan is great for exploring, but some features are limited or locked. Here's exactly what changes when you upgrade.

Dataset card tabs

On Free, only two tabs are fully readable. The rest show a small lock and a blurred preview.

DescriptionMetadataAnalysisBusiness ApplicationsHow to enrichResponsible AIQ&A

Other Free-plan limits

  • Results — you see the top 3 plus the next 10 (further results are hidden)
  • Saved datasets — up to 10 bookmarks
  • Workspace — basic, with some tabs hidden
  • Query length — up to 140 characters per search
  • Advanced metadata — preview only
Upgrading to Startup or Pro unlocks every dataset card tab, more results, the advanced workspace, and longer queries.
08

Profile & settings

Manage your account from Settings in the left menu. This is where you keep your profile and account preferences up to date.

  • Profile information — your name and account details.
  • Account preferences — your general settings for the app.
  • Subscription & billing — your plan, credits, and payment details live under Subscription Plan in the left menu, or via the credit-card icon in the top bar (see Subscriptions & billing).
Your current plan and credit balance are always shown in the top-right corner, so you can check where you stand without opening Settings.
08.1

Notifications

The bell icon in the top bar keeps you up to date, so you don't have to keep checking back manually. Notifications cover things like:

  • Custom dataset order updates — when an order moves to QA Review, is Completed and ready to download, or is Rejected.
  • Account activity — relevant changes to your account.

Select the bell to see your recent notifications at any time.

Order updates also show in My Datasets Orders, where you can hover a status for the estimated delivery date.
08.2

Getting support

Need a hand? A few options:

Contact Sales

For plan questions, higher-volume needs, or enterprise requirements, reach our team from Contact Sales in the app.

Help Center

You're here. Browse the categories on the left for step-by-step guidance on every part of the workspace.

On Pro, if you need a higher research allowance than your plan includes, contact us — we can discuss higher-volume options.