codelucas/ceo-quotes-verified-sample
ποΈ CEO Transcripts β Verified Executive Interviews The World's Largest Database of Verified C-Suite Transcripts 20,000+ Executives Β· 100,000+ Transcripts Β· 400,000+ Quotes Β· S&P 500 + NASDAQ + Global Leaders π₯ What's In This Sample? This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy. Executive Role Why They Matter Jensen Huang CEO, NVIDIA Every AIβ¦ See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.
<div align="center">
ποΈ CEO Transcripts β Verified Executive Interviews
The World's Largest Database of Verified C-Suite Transcripts
  [](mailto:lucas@ceointerviews.ai)
20,000+ Executives Β· 100,000+ Transcripts Β· 400,000+ Quotes Β· S&P 500 + NASDAQ + Global Leaders
</div>
π₯ What's In This Sample?
This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy.
Sample Date Range: 3 months (August β Nov 2025) Full Dataset: 2007-Present (updated daily)
π‘ The Problem We Solve
Alpha lives in unguarded moments.
The media playbook for executives has fundamentally shifted:
The signal is there. But it's buried in 10,000+ hours of fragmented audio.
CEOInterviews.ai transforms this chaos into structured, queryable, backtestable data.
π― Example: Powell vs Dimon on Recession Risk
Research Question: "What did Jerome Powell and Jamie Dimon say about recession risk in 2022?"
Using Our API:
import requests
API = "https://ceointerviews.ai/api"
headers = {"X-API-Key": "your_api_key"}
powell_quotes = requests.get(f"{API}/quotes/", params={"entity_id": 15847, "keyword": "recession"}, headers=headers).json()
dimon_quotes = requests.get(f"{API}/quotes/", params={"entity_id": 12903, "keyword": "recession"}, headers=headers).json()
for q in powell_quotes["results"][:3]:
print(f"[Powell {q['source_created_at'][:10]}] {q['text'][:150]}...")
for q in dimon_quotes["results"][:3]:
print(f"[Dimon {q['source_created_at'][:10]}] {q['text'][:150]}...")Sample Output:
[Powell 2022-06-15] "We're not trying to induce a recession now, let's be clear about that.
We're trying to achieve 2% inflation..."
[Dimon 2022-06-01] "You know, I said there's storm clouds but I'm going to change it...
it's a hurricane. Right now, it's kind of sunny, things are doing fine..."
[Powell 2022-09-21] "We have got to get inflation behind us. I wish there were a painless
way to do that. There isn't..."
[Dimon 2022-09-26] "This is serious stuff... it's a different environment than we've ever
seen before... the Fed has to meet this now..."This is the alpha. Dimon called the "hurricane" 2 weeks before Powell acknowledged pain was coming.
π Dataset Schema
Each row is a notable quote with full context about the executive and source video.
β What Makes CEOInterviews Different?
1. Appearance Date Detection
Most datasets only have publish_date. But a podcast uploaded today might contain an interview from 6 months ago.
We detect when the executive actually spoke.
# Example: Interview recorded in January, published in March
{
"publish_date": "2022-03-15", # When YouTube uploaded
"appearance_date": "2022-01-20", # When Buffett actually spoke β
}This is critical for backtesting. You need to know when the market could have known, not when the video appeared.
2. AI + Human Verification
Every transcript passes through:
- π€ Thumbnail Analysis: Is the executive actually in the video?
- π€ Transcript Verification: Is this their voice, not dubbed/AI-generated?
- π€ Quality Scoring: Is the transcript complete and accurate?
No deepfakes. No dubbing. No secondhand reporting.
3. Structured Quote Extraction
We don't just give you transcriptsβwe extract the market-moving moments:
{
"text": "We're seeing something we haven't seen in 40 years...",
"is_notable": True,
"is_financial_policy": True,
"topics": ["inflation", "monetary_policy", "fed"],
"timestamp_in_video": "00:23:45"
}π Use Cases
π Sample vs Full Dataset
π Get Full Access
This sample is <1% of our full dataset.
Full Dataset Includes:
- β Every S&P 500 and NASDAQ CEO
- β Global political leaders (Presidents, Prime Ministers, Fed chairs)
- β Top AI founders (Altman, Hassabis, etc.)
- β Legendary investors (Buffett, Dalio, Ackman, Fink)
- β Daily updates
- β RESTful API with full pagination
- β CSV/JSON export for model training
- β White-glove enterprise support
Pricing
Contact
π§ Email: lucas@ceointerviews.ai π Website: ceointerviews.ai π API Docs: ceointerviews.ai/api_docs
π License
CC-BY-NC-ND-4.0 (Creative Commons Attribution-NonCommercial-NoDerivatives 4.0)
What This Means:
For commercial use, contact [lucas@ceointerviews.ai](mailto:lucas@ceointerviews.ai)
π οΈ Built By
<div align="center">
Lucas Ou-Yang Former Engineering Manager and Staff ML Engineer @Coinbase, @Tiktok, @Meta superintelligence labs
Building institutional-grade datasets for quantitative research.

</div>
π£ Citation
If you use this dataset in research, please cite:
@misc{ceointerviews2024,
author = {Ou-Yang, Lucas},
title = {CEOInterviews.ai: Verified Executive Transcript Dataset},
year = {2024},
publisher = {HuggingFace},
url = {https://huggingface.co/datasets/codelucas/ceo-transcripts-verified-sample}
}<div align="center">
β Star this dataset if you find it useful!
Questions? Reach out at lucas@ceointerviews.ai
</div>
