Pricing · Dataset Finder

Simple pricing for AI teams and solo builders.

Search a hand-curated catalogue of vetted AI datasets for free. Upgrade when you need deeper dataset intelligence, unlimited workspace projects, or custom data credits.

Free forever plan · No credit card required · Cancel anytime

Trusted by leading AI builders
Tier 01 / Free
Free

For individuals exploring AI datasets.

€0/month
No credit card required
Includes:
  • Monthly research allowance to explore the catalog
  • Full dataset discovery across tens of thousands of curated datasets
  • Top recommendations + a curated set of dataset results
  • Dataset card essentials
  • Save a handful of favorite datasets
  • 1 demo project workspace
  • Order custom datasets with credits
Start Free
Tier 03 / Pro
Pro

For advanced AI teams and professionals.

€29.90/month
Everything in Startup, plus:
  • High-volume monthly research for power users
  • Unlimited workspace projects with project linking
  • Full AI projects & recipes library · community recipes reviewed & validated by our team
  • Export & documentation · dataset reports, compliance reports
  • Training data inventory · lineage, audit trail, GDPR-ready
  • 25% discount on custom dataset credits
  • Priority support
Go Pro

All plans include access to custom dataset ordering with platform credits. Custom data is billed separately based on volume, modality, and complexity. Need higher search volume or enterprise features? Talk to Sales →

Tier 04 / Enterprise

Large-scale programmes, quoted to fit.

Past a few hundred thousand items, or working under procurement, security review or data-residency constraints, the self-serve plans stop being the right shape. Tell us what you are building and we will scope it and come back with a quote.

Typical reply within one business day · No commitment

Everything in Pro, plus
  • A real quote, not credits · scoped and priced by Innovatiana, from years of running large-scale data annotation programmes
  • Custom SLAs & NDAs with a security and compliance review
  • On-premise delivery for regulated or residency-bound data
  • Dedicated labelers embedded in your team and your tools
  • Audit-ready governance · lineage, audit trail, GDPR & EU AI Act

What you get with every plan

01

Search what already exists

Describe the model you are building in plain language. We surface datasets our team has already reviewed, scored and documented, so you skip weeks of hunting through directories.

02

Order what doesn't

When no dataset fits, commission one. Pay with platform credits, our trained labelers and domain experts produce it end-to-end, and you download the result. No procurement cycle.

03

Keep track of all of it

Projects, collections, notes and lineage in one place, so your team stops losing decisions in Slack threads and always knows which data trained which model.

§ 02 · The Difference
Why dataset finder

More than a dataset directory.

The status quoDataset Finder
Browse 5 different directories, filter by keywordDataset discovery across tens of thousands of curated AI datasets
Read research papers to evaluate dataset qualityAutomated analysis on 30+ criteria per dataset
Hire labelers, write guidelines, run QA yourselfCustom data on-demand · labelers, QA, delivery handled
Quote-and-wait procurement (weeks to start)Self-serve credit ordering · kick off in minutes
Bring your own pipeline, scale headcount internallyPlug our trained annotators into your existing tools
Decisions live in Slack threads & lost wikisProject workspace with notes, links, and lineage
Reinvent the wheel for every fine-tuning projectCommunity AI recipes reviewed & validated by our team
Scramble before audits, scattered evidenceBuilt for GDPR & EU AI Act audit-ready by default
Discovery, labeling, governance across 4+ different vendorsOne workspace for the full training data lifecycle
§ 03 · Customers
Tested and approved

What our customers say.

Dataset Finder is built by the Innovatiana team. Here is what the AI teams we work with say about the data we deliver.

Innovatiana helps us carry out data labeling tasks for our classification and text recognition models, which requires a careful review of thousands of real estate ads in French. The work provided is of high quality and the team is stable over time. The deadlines are clear, as is the level of communication.
Tim KeynesChief Technology Officer, Fluximmo
Several Data Labelers from the Innovatiana team are integrated full time into my team of surgeons and Data Scientists. I appreciate the technicality of the Innovatiana team, which provides me with a team of medical students to help me prepare quality data, required to train my AI models.
Dan D.Data Scientist & Neurosurgeon, Children's National
Working with Innovatiana has been a great experience. Their team was reactive, rigorous and very involved in our project to annotate and categorize industrial environments. The quality of the deliverables was there, with real attention paid to the consistency of the labels.
Kasper LauridsenAI & Data Consultant, Solteq Utility Consulting
Innovatiana helps us a lot in reviewing our data sets in order to train our machine learning algorithms. The team is dedicated, reliable and always looking for solutions. I also appreciate the local dimension of the model, which allows me to communicate with people who understand my needs.
Henri RionCo-Founder, Renewind
Innovatiana embodies exactly what we want to promote in the data annotation ecosystem: an expert, rigorous and resolutely ethical approach. Their ability to train and supervise highly qualified annotators, while ensuring fair and transparent working conditions, makes them a model of their kind.
Bill HeffelfingerCEO, CVAT (2023–2024)
Innovatiana is deeply committed to ethical AI. The company ensures that its annotators work in fair and respectful conditions, in a healthy and caring environment. Innovatiana applies fair working practices for Data Labelers, and this is reflected in terms of quality.
Sumit SinghProduct Manager, Labellerr
§ 04 · FAQ

Questions, answered.

The things AI teams ask us most, including exactly how the data works.

Datasets & the catalogue
What is Dataset Finder?

Dataset Finder is a workspace for AI training data. You search a curated catalogue of tens of thousands of machine learning datasets in plain language, read a full dataset card before committing, organise what you find against your AI projects, and commission custom data annotation when nothing suitable exists. It replaces the scattered mix of directory browsing, paper reading, spreadsheets and labeling vendors most teams juggle today.

Do you host or store the datasets?

No. Dataset Finder does not host, store, or redistribute datasets. We are a curated catalogue and a workspace built on top of data that already exists in the world, not a file repository.

What we maintain is the metadata and the intelligence around each dataset: where it lives, who published it, its licence, size and format, quality scores, biases, limitations, and suitable use cases. When you bookmark a dataset, your workspace stores that record and your notes, never the files. Dataset Finder sends you to the origin, whether that's Hugging Face, an academic page, a government portal or your own storage, where you access it under its own licence terms.

The one exception is custom data: when you commission an annotation project, our team produces and delivers that dataset to you directly.

Is it legal? What about dataset licences?

Yes, because we never take possession of or resell anyone's data. You obtain each dataset from its original publisher, under that publisher's licence. To make that easier, every dataset card carries a Licence field and a Licence check, so you can see whether a dataset is permissive, non-commercial or research-only before you invest time in it. We surface the information; you keep the relationship with the source. Always review the licence yourself before training on a dataset.

How is this different from browsing Hugging Face or Kaggle?

Those are repositories: enormous, keyword-driven and largely unevaluated. You still judge quality yourself, usually by reading papers and sampling files. Dataset Finder searches across those sources and adds the layer they don't have: human curation, quality and bias analysis, label-density and class-balance signals, suitability scoring, and a workspace to track decisions. And when nothing in any repository fits, you can order the annotated data instead of giving up.

What does “curated” actually mean here?

It means a human looked at it. Datasets are reviewed and enriched by our team rather than crawled in bulk, and we refresh the catalogue continuously: new datasets, corrected metadata, updated licences, dead and moved links swept out. That's the difference between a search index and a catalogue you can trust.

Can my whole team work in it together?

Yes. Create a workspace per team, client or product line and invite colleagues into it. Bookmarked datasets, statuses, notes, AI projects and cloned recipes are shared, so two people don't evaluate the same dataset twice and the work survives someone changing role.

Custom data & annotation
What if the dataset I need doesn't exist?

Order it. Describe the data you need, attach your annotation guidelines and a few labeled examples, and our team handles data collection, labeling, quality assurance and delivery. You can self-serve with credits for typical projects up to a few hundred thousand items, or talk to our team for large-scale work with custom SLAs, NDAs or on-premise delivery.

What types of data annotation can you produce?

The full range of supervised labeling work, across image, video, text, audio, multimodal and document data:

Computer vision. Bounding boxes, polygons, semantic and instance segmentation masks, keypoints and landmarks, object tracking across video frames, image classification and tagging.
NLP & documents. Named entity recognition (NER), text classification, sentiment and intent labeling, relation extraction, OCR correction, document layout and key-value extraction, translation and multilingual annotation.
Gen-AI & LLMs. Instruction and prompt/response pairs for supervised fine-tuning (SFT), preference ranking and RLHF, red-teaming, content moderation, factuality and hallucination review.
Audio. Transcription, speaker diarization, event and intent labeling.

We also cover regulated domains such as medical imaging and clinical text, where annotation is carried out by domain experts rather than generalists.

Can you work inside our existing labeling infrastructure?

Yes, and this is usually the preferred setup. Our annotators can log into your annotation platform and work in your environment, your schema and your review workflow, so the labeled data lands where your pipeline already expects it and never leaves your tenancy. We're experienced across the common tools, including Label Studio, CVAT, V7, Labelbox, Encord, Roboflow, SuperAnnotate and Prodigy, and can work in an in-house or proprietary tool too.

If you don't have a platform, we'll provide and administer a secure one for your project, so tooling never becomes the blocker.

How is this different from Mechanical Turk or crowdsourcing?

Fundamentally different model. Crowdsourcing marketplaces distribute micro-tasks to an anonymous, rotating pool paid per click, which is why they struggle with complex guidelines, domain knowledge, consistency over time and any kind of traceability.

We don't use a crowd. Annotation is done by in-house teams who are recruited, trained and fairly employed by Innovatiana, working from our own delivery centre, supervised by a dedicated Data Labeling Manager who is your point of contact. The same people stay on your project, so guidelines compound instead of resetting, and every annotation is traceable to a known annotator. For specialist work we staff domain experts such as medical students and clinicians, linguists, lawyers and developers, rather than generalists.

Practically, that means higher inter-annotator agreement, workable complex taxonomies, and an audit trail you can show a regulator. It's also why we can commit to ethical sourcing rather than opaque subcontracting.

How do you guarantee annotation quality?

Quality assurance is built into the workflow, not bolted on at the end. We start from your guidelines and a calibration round on a sample, then run manual sampling reviews, inter-annotator agreement (IAA) measurement, gold-standard comparison and automated consistency checks throughout the project. A dedicated manager monitors throughput and label consistency, and edge cases are escalated back to you rather than guessed. You get corrected, production-grade ground truth, not a first pass.

What delivery formats do you support, and how fast?

We deliver in the format your training pipeline expects: COCO, YOLO, Pascal VOC, JSON, JSONL, CSV, XML, or your own custom schema, transferred securely or written straight into your storage. On timing, we assess your needs and run a test batch where useful within about 48 hours, so you can validate quality on real examples before committing to full volume. Delivery timelines then depend on volume and annotation complexity, and are agreed up front.

Is our data kept confidential and compliant?

Yes. We assess the sensitivity of whatever you entrust to us and apply the matching controls: NDAs, restricted access, secure transfer, and working inside your environment where required. Our practice is built around GDPR and EU AI Act readiness, with full traceability of who annotated what, which matters when you have to document your training data provenance. Data handling terms are agreed before a project starts.

Plans & company
How does pricing work? Is there really a free plan?

There is, and it needs no credit card. Free gives you dataset discovery across the whole catalogue, dataset card essentials and one demo project. Startup (€9.90/mo) unlocks full dataset cards, unlimited saved datasets and multiple projects. Pro (€29.90/mo) adds unlimited projects, the full recipes library, exports and compliance reports, and training data lineage. Custom annotation is priced separately with credits, charged per task or per dataset delivered with no subscription or set-up fee, and paid plans carry a discount on those credits.

Who is behind Dataset Finder?

Dataset Finder is built by Innovatiana, a data labeling company with in-house teams of professional annotators and AI trainers who have delivered training data to 200+ AI builders since 2021. Not a crowdsourcing marketplace: fairly employed annotators, full traceability, and GDPR and EU AI Act readiness by default.

Still have a question? Talk to our team and we'll reply within one business day.

Get started today

AI models improve.
Your training data workflow
should too.

Discover tens of thousands of curated AI datasets, order custom training data, and turn fragmented data pipelines into production-ready AI assets, all in one workspace.

FAQ