Leaderboard 8 AI models vs. a 1908 Urdu text: the best got 1 word in 11 wrong

Expertise AI can't scrape. Verified before it trains.

The vetted expert workforce for frontier AI labs and AI-data vendors. Physicians, lawyers, scientists, engineers and native speakers of 100+ languages, each identity- and credential-checked before they write, rank or red-team a single item of your data.

For AI teams Request a pilot
For experts Apply as an expert
551
open roles
142
languages hiring
33
fields hiring
  • Clinical medicine
  • Pharmacology & drug safety
  • Law & regulation
  • Physics & chemistry
  • Aerospace & space
  • Software & mathematics
  • Defence & security
  • Finance & accounting
  • Agronomy
  • Education
  • Art & design
  • Tibb & traditional medicine
  • العربية Arabic
  • हिन्दी Hindi
  • বাংলা Bengali
  • اردو Urdu
  • Kiswahili
  • Bahasa Indonesia
  • Hausa
  • Tagalog
  • Vietnamese
  • English

Two ways in

One network of verified experts. Two sides of the same trust.

AI teams get data from people whose identity and credentials are on record. Experts get fair, paid work that uses what they actually know.

For AI labs & data vendors

Expert data you can audit, item by item.

  • RLHF, preference data and expert rewrites
  • Held-out evaluations and red-teaming
  • RL environments and verifiable tasks
  • Specialists staffed onto your own projects

For experts

Paid, remote AI work in your own field.

  • Rate agreed per project, before you start
  • Paid in US dollars by bank transfer, Payoneer or Wise
  • A written contract and consent over how your work is used
  • No fee to apply, ever

Proof, published

We gave AI models a 1908 Urdu text. The best still got 1 word in 11 wrong.

A vetted native expert typed the truth blind, every disagreement was checked against the scan, and every word was scored. On our public leaderboard the top model, Claude Fable 5.1, reproduced 91.08% of words exactly and still made a mistake on 111 of 161 lines. Many errors carry no warning flag, so nothing in the output tells a reviewer to look twice.

Line with a mistake, best model · 111 Read correctly · 50
91.08%
best strict word accuracy (Claude Fable 5.1)
111/161
lines with a mistake, even for the best model
2,378
words and marks typed blind by a vetted native expert
8
models ranked so far, more added as they finish

The same method runs in any field: truth written blind by a vetted expert, adjudicated against the source, scored item by item.

What we deliver

Data that makes models measurably better at the hard parts.

Managed pods of verified experts, scoped to your capability gap and quality bar.

Expert RLHF & preference data

Rankings, rewrites and written rationales from credentialed professionals, not generalists guessing.

Evaluations & red-teaming

Gradable test sets and adversarial probes in medicine, law, science, engineering, finance, defence and 100+ languages.

RL environments & verifiable tasks

Tasks with checkable answers and reward signals, designed by the people who do the work.

Reasoning in 100+ languages

Native speakers with subject expertise, from major world languages to low-resource ones, not machine translation of English prompts.

Provenance-licensed datasets

Consented, documented and exclusive-ready, with a chain of custody for every item. Built for a legal landscape where provenance decides what can be used.

How verification works

Five checks before anyone touches your data.

Every expert passes all five before client work, and the result of each check is kept on record, so you can see who produced your data and why they were qualified to.

  1. Identity

    Government-ID verification with consent, a liveness check and duplicate-account detection.

    On record: ID match

  2. Credentials

    Degrees, licenses and registrations checked with the issuing body: medical boards, bar associations, engineering institutions, universities.

    On record: registry check

  3. Domain assessment

    A role-specific test, written and graded by senior practitioners in the field.

    On record: score vs. pass mark

  4. Paid calibration

    Trial tasks scored against gold answers before any client work begins.

    On record: agreement with gold

  5. Continuous quality

    Ongoing agreement scoring and spot audits. Anyone below the bar comes off the project.

    On record: per-batch audit

Provenance by default

  • Consent recorded for every contributor
  • Chain of custody for every item
  • Community-consent protocol for cultural material, following the CARE principles
  • Exportable audit trail with each delivery

Security roadmap

  • SOC 2 Planned
  • ISO/IEC 27001 Planned
  • ISO/IEC 42001 Planned

Designed for least-privilege access, encrypted storage and per-project isolation from day one. We only claim a certification once an auditor has issued it.

Fair work

  • Written contracts for every contributor
  • Rate agreed per project
  • Payouts on a published schedule
  • No fee to apply, ever

For AI teams

From capability gap to delivered data in four steps.

  1. Step 01

    Scope

    Tell us the capability gap, the domain and your quality bar. We design the task and the rubric with you.

  2. Step 02

    Match

    We assemble a pod of verified experts whose credentials fit the task, not whoever is online.

  3. Step 03

    Produce

    Dual review, gold tasks and calibration on every batch, with agreement scores you can see.

  4. Step 04

    Deliver

    Data in your format, with its provenance ledger attached.

Domains

Every field a frontier model is asked about. Experts who can check it.

We recruit and vet experts across the professions, the sciences and the world's languages: wherever a model's answer needs a qualified human to write it, grade it or break it.

Medicine, pharma & Tibb

Clinical reasoning, diagnostics, pharmacology and drug safety, plus evidence review of Tibb-e-Unani and other traditional medicine.

Law & policy

Statutory and case-law reasoning, contracts, IP and patents, and compliance, across common-law, civil-law and religious systems.

Sciences

Physics, chemistry, biology and mathematics: proofs, derivations and reasoning rubrics written by researchers.

Engineering & software

Civil, electrical, mechanical and software engineering; code review, systems and technical protocols.

Aerospace & space

Aeronautics, propulsion, flight systems, avionics and space operations.

Finance & business

Banking, capital markets, accounting, tax, insurance, strategy and operations, including Islamic finance.

Defence & security

Strategic analysis, cybersecurity, OSINT and adversarial red-teaming of model behavior.

Education & academia

Teachers, professors and assessment writers for curricula, exam items and step-by-step explanations.

Art & design

Visual art, design, music, film and architecture, judged by working practitioners.

Languages

100+ languages: native-speaker reasoning, translation QA, transcription and dialect coverage, from major world languages to Urdu and other regional languages.

Agriculture & climate

Agronomy, crop and soil science, irrigation, veterinary science and climate adaptation.

Heritage & humanities

History, philosophy, religious and literary traditions, manuscripts and archives, documented and consented.

Vettedlance research

Measuring what frontier models don't know yet.

The Urdu study is published. Three more expert-built benchmarks are in development, and each will publish after independent expert review, with methodology and agreement scores.

VL-LawIn development

Statutory reasoning across jurisdictions

Statute and case-law questions graded against verifiable citations, jurisdiction by jurisdiction.

Format
Questions with verifiable citations
Written by
Practising lawyers
Scoring
Per item, citation-checked
VL-LangIn development

Low-resource language reasoning

Instruction-following and reasoning in languages with little presence in training data, written by native speakers.

Format
Native-written prompts and reference answers
Written by
Native speakers with subject expertise
Scoring
Blind, adjudicated
VL-ClinicIn development

Clinical & pharmacology safety

Clinical reasoning and traditional-medicine claims graded against evidence and flagged for safety.

Format
Clinical and drug-safety cases
Written by
Physicians and pharmacists
Scoring
Evidence-graded, safety-flagged

Why it matters beyond Urdu: frontier models score up to 24.3 points lower when the same questions are asked in a low-resource language (MMLU-ProX, 2025), and 88% of the world's languages were found to be left behind by language technology (ACL, 2020).

Why Vettedlance

Built for trust, not volume.

How Vettedlance compares with generic crowd platforms
CriterionGeneric crowd platformsVettedlance
Who does the workWhoever is onlinePractitioners matched to the task by field and language
IdentitySelf-reportedGovernment-ID verification with consent
CredentialsCV claimsChecked with the issuing body
Quality controlSpot checks after deliveryGold tasks, calibration and agreement scores on every batch
ProvenanceRarely documentedA record for every item: who, when, which guideline, how reviewed
CoverageGeneralists, English-firstRecruited by domain, in 100+ languages
Contributor termsOpaque, variableWritten contract, rate agreed before work starts

Open roles

551 ways to put your expertise to work. Paid in dollars.

Physicians, lawyers, scientists, engineers, analysts, artists and native speakers, across 33 fields and 142 languages. Apply once and get matched to paid, remote AI projects.

Physician (MBBS/FCPS) for Clinical AI Evaluation

Paid in USD

Medicine & health

Expert · Talent pool

Lawyer (Advocate) for Legal Reasoning Evaluation

Paid in USD

Law & policy

Expert · Talent pool

Software Engineer for AI Code Evaluation

Paid in USD

Engineering & software

Expert · Talent pool

Mathematician (MS/MPhil/PhD)

Paid in USD

Sciences

Expert · Talent pool

Unani Medicine Physician (BEMS/MD Unani)

Paid in USD

Traditional & Eastern medicine

Expert · Talent pool

Mufti / Fatwa Research Scholar (Iftaa)

Paid in USD

Religion, Sufism & folklore

Expert · Talent pool

Aerospace Engineer for AI Evaluation

Paid in USD

Aerospace & space

New opportunity

Clinical Pharmacologist for RLHF

Paid in USD

Pharma & life sciences

New opportunity

Defence Analyst for AI Red-Teaming

Paid in USD

Defence & security

New opportunity

Urdu Transcriptionist, Historical Texts

Hiring now

Paid in USD

South Asian languages · اردو

Transcription · Open now

Swahili Language Expert for AI Training & Evaluation

Paid in USD

African languages · Kiswahili

Expert · Talent pool

Arabic (Modern Standard) Language Expert for AI Training & Evaluation

Paid in USD

Middle Eastern & Central Asian languages · العربية

Expert · Talent pool

For experts

Your expertise, paid in dollars. On time.

If you hold a professional license, an advanced degree, or native fluency in a language models handle badly, your knowledge is exactly what AI systems are missing.

  • Rate agreed per project, before you start
  • Paid in US dollars by bank transfer, Payoneer or Wise
  • A written contract and consent over how your work is used
  • Flexible, remote work that fits around your practice
  • No fee to apply, ever
Apply as an expert
  1. 01

    Apply

    About ten minutes. Tell us your field, credentials and languages.

  2. 02

    Verify

    Confirm your identity and credentials with your consent.

  3. 03

    Assess

    A short test written by senior people in your field.

  4. 04

    Calibrate

    Paid trial tasks so you learn the quality bar.

  5. 05

    Work

    Join projects that match your expertise.

FAQ

Questions, answered plainly.

Are you already working with AI labs?

We're onboarding our founding expert cohort and building our first datasets and benchmarks. We don't list clients or partners we haven't signed, and we never will.

Can we start with a pilot?

Yes. Most teams start with a fixed-price pilot scoped to one capability gap, then scale once the quality is proven on their own evaluation.

How do you verify experts?

Five checks: government-ID verification with consent, credential checks with the issuing body, a domain assessment written by senior practitioners, paid calibration tasks, and continuous quality scoring. Nobody touches client work until all five are complete.

Do you work with AI-data vendors as well as labs?

Yes. We can staff verified specialists onto a vendor's own projects, or deliver finished data and evaluations, with the same verification and provenance records.

How do you handle copyrighted or cultural material?

We only use material we have the right to use, and we document where every item came from. For cultural and community knowledge, we follow the CARE principles for Indigenous data governance and record consent before anything is collected.

How are experts paid?

In US dollars, at a rate agreed before a project starts, by bank transfer, Payoneer or Wise, on a published schedule. Every contributor signs a written contract.

Which domains and languages do you cover?

We staff by demand rather than a fixed list: medicine and pharma, law, the sciences, engineering including aerospace, software, finance, defence, education and the arts, plus English, the major world languages and the low-resource languages where native-speaking experts are hardest to find. We add a language once we can verify enough native speakers with subject expertise to meet our quality bar.

Is my data secure?

We design for least-privilege access, encrypted storage and per-project isolation from day one. SOC 2, ISO/IEC 27001 and ISO/IEC 42001 are on our roadmap, and we only claim a certification after an auditor issues it.

Get started

Put verified expertise behind your model.

Start with a scoped pilot, or join the expert network. Either way, verification comes first.

For AI labs & data vendors

Scope a pilot in your domain, languages and quality bar. We reply within two business days.

For experts

Paid, remote AI work in your field. Free to apply, with a written contract for every project.