alvaroluceroo/llm-api-pricing
LLM API pricing dataset Prices of the current large language model APIs, as published by their makers, with the day each price was last verified and the maker's page it was read from, plus a record of every change to those prices. This is the data behind the LLM API pricing table of AI Signal, published here as files so it can be versioned, diffed and cited. The same files are served at aisignalhq.com/data/ and versioned on GitHub at alvaroluceroo/llm-api-pricing-dataset. This… See the full description on the dataset page: https://huggingface.co/datasets/alvaroluceroo/llm-api-pricing.
LLM API pricing dataset
Prices of the current large language model APIs, as published by their makers, with the day each price was last verified and the maker's page it was read from, plus a record of every change to those prices. This is the data behind the LLM API pricing table of AI Signal, published here as files so it can be versioned, diffed and cited.
The same files are served at aisignalhq.com/data/ and versioned on GitHub at alvaroluceroo/llm-api-pricing-dataset. This dataset is a mirror of them: every commit is a snapshot of what the site published that day.
from datasets import load_dataset
models = load_dataset("alvaroluceroo/llm-api-pricing", "models", split="train")
changes = load_dataset("alvaroluceroo/llm-api-pricing", "price-changes", split="train")The models configuration reads the rows of models.json, with every long-prompt tier and the promotion as nested fields; price-changes reads the rows of price-changes.json. models.csv holds the same model rows flattened, for spreadsheets and pandas (below). Both configurations use the JSON files because a Hugging Face dataset loads all its configurations with one reader.
Files
Both JSON files share one envelope: name, publisher, homepage, generatedAt (when that copy was built, UTC), license, licenseUrl, attribution, rowCount and rows.
models.csv starts with three lines beginning with # that carry the name, build time, licence and attribution, followed by a header line. Skip them with pandas.read_csv(path, comment="#") or the equivalent option of your reader; no value in the file contains #.
import pandas as pd
url = "https://huggingface.co/datasets/alvaroluceroo/llm-api-pricing/resolve/main/models.csv"
df = pd.read_csv(url, comment="#")Scope
The model files hold models whose status is current, charged per token, with text output, and priced by their maker. Deprecated and retired models are not included. A row whose maker was checked and publishes no per-token price keeps its prices empty and says no-published-price in priceStatus: no figure in these files is estimated, rounded or filled in.
Left out on purpose: prices on inference hosts (Together, Fireworks, DeepInfra, ...), second price lists such as Batch or Priority, and the site's internal notes and working fields.
Fields of models.json
models.csv has the same fields with nested ones flattened: the first long-prompt tier goes into longTierCondition, longTierInputPerMTok and longTierOutputPerMTok, and longTierCount says how many tiers the row has (the others are only in the JSON); the promotion goes into promoEndsOn and promoAtLeastUntil. Columns, in order: id, apiId, maker, makerName, name, status, priceStatus, inputPerMTok, outputPerMTok, cachedInputPerMTok, cacheWritePerMTok, cacheWrite1hPerMTok, priceRegion, longTierCondition, longTierInputPerMTok, longTierOutputPerMTok, longTierCount, contextWindow, maxOutput, promoEndsOn, promoAtLeastUntil, verifiedAt, sourceUrl, pageUrl.
Fields of price-changes.json
How each price is verified
Every figure is read on a page the maker publishes (its pricing page, a model page or its documentation), and the row records that page in sourceUrl and the day in verifiedAt. In short:
- Weekly detector. Every Monday a script downloads the page each row cites for the makers it can read and checks that the row's input and output prices appear on it. A match must be backed by a line quoted verbatim from the page; a mismatch, a changed layout or a missing row goes to the editor, who reads the page and corrects the row by hand.
- Manual review. Makers whose pages a script cannot read, and rows with no published price, are checked by hand in a weekly review.
- Corrections are recorded, never silent. Every change, whether by a provider or a correction of the site's own reading, is an entry in
price-changes.jsonwith its date and source.
The full process, and what it does not guarantee, is described in the methodology.
Providers change prices without notice. The authoritative source for every price is the maker's page in sourceUrl: confirm there before committing spend.
Update frequency
The files are regenerated on every build of aisignalhq.com, and a new commit is pushed here and to the GitHub repository automatically whenever the data changes. A build that changes nothing but the build time (generatedAt) does not produce a commit. verifiedAt on each row, not the date of the commit or generatedAt, says when its price was last confirmed.
Licence
The data is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) licence. You may copy, share and adapt it, including commercially, provided you give attribution. The full legal text is at creativecommons.org/licenses/by/4.0/legalcode.
Attribution
Attribution line, ready to copy, for a page, an app or a chart:
LLM API pricing data by AI Signal (https://aisignalhq.com/api-pricing/), licensed under CC BY 4.0.Link to aisignalhq.com and say whether you changed the data. For a single price, citing the maker's page in sourceUrl as well is the most checkable form. For papers, use the CITATION.cff of the GitHub repository.
Errors
To report an error, open a discussion here, an issue on GitHub, or write to the editor from the About page.
