Product lookup
Find what published sources say about a product from its GTIN, ASIN, part number or page address, with the source of every fact.
Product lookup takes whatever you know about a product (a barcode, an ASIN, a part number, a page address) and returns what published sources say about it. Every fact comes back with the source that states it, so you can show where it came from and decide what to trust.
Use it to fill gaps in a catalogue, to check a supplier's data against the open record, or to gather facts before you generate content.
Look up one product
curl https://ace.authoritas.com/api/v1/enrichment/lookup \
-H "Authorization: Bearer $ACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "q": "4006381333931", "market": "GB" }'q is a bare value. Its type is worked out for you: a barcode, an ISBN, an ASIN, a page address or a part number. If you already know what you have, name it:
{
"identifiers": { "gtin": "4006381333931" },
"hints": { "brand": "STABILO" },
"market": "GB"
}Identifiers
| Field | What it is | Where it is looked up |
|---|---|---|
gtin (or ean, upc) | GTIN-8, UPC, EAN-13 or GTIN-14 | Every source |
isbn | ISBN-10 or ISBN-13 | Every source |
asin | Amazon identifier | Your connected Amazon store and your catalogue |
mpn | Manufacturer part number | Your catalogue. With hints.brand, the web too |
sku | Your own stock code | Your catalogue only |
url | A product page | That page is read |
Barcodes are checked against their check digit, so a mistyped number is reported instead of being searched for. Every written form of one number is treated as the same product: 036000291452, 0036000291452 and 00036000291452 are one barcode.
What comes back
{
"data": {
"operation": "lookup",
"items": [
{
"index": 0,
"status": "ok",
"warnings": [],
"match": { "status": "matched", "confidence": 0.95, "identifiers": { "gtin": "04006381333931" } },
"product": { "title": "…", "brand": "…", "attributes": { "color": { "value": "Azure" } } },
"facts": [
{
"field": "attributes", "name": "color", "value": "Azure",
"source": "web_search", "url": "https://…", "retrievedAt": "2026-10-09T09:12:00Z",
"matchMethod": "exact_gtin", "extraction": "structured", "confidence": 0.95
}
],
"sources": [
{ "id": "open_data", "status": "miss" },
{ "id": "web_search", "status": "hit", "urls": ["https://…"] }
],
"conflicts": [],
"usage": { "searchCalls": 1, "cacheHits": 0 }
}
],
"summary": { "total": 1, "matched": 1, "ambiguous": 0, "notFound": 0, "errors": 0, "searchCalls": 1, "cacheHits": 0 },
"attributions": []
}
}matchsays whether the identifier was matched and how sure that is.matchedmeans the sources agree on one product.ambiguousmeans they found more than one, or only weak evidence, andmatch.candidateslists what they found.not_foundis a normal answer and still returns200.productis the best value per field across sources.factsis every statement, with the source and page that made it, when it was read, and a confidence.sourcessays what each source did. A source you cannot use reportsunavailablewith a reason code.conflictslists where two sources disagree, with the value that was chosen.
Use include to return only the sections you want, for example "include": ["product"].
Confidence
A fact's confidence combines how the product was matched with how the value was read.
| How it was found | Confidence |
|---|---|
| A record keyed on the identifier itself (your catalogue, Amazon, Open Food Facts) | 0.95 to 0.99 |
| A page whose own structured data states the identifier | 0.95 |
| A page that mentions the identifier | 0.8 |
| A page a search cited that could not be read to confirm | 0.5 |
Facts below minConfidence (default 0.6) are left out of product. They are still listed under facts, so you can see them and lower the threshold if you choose.
Sources
| Source | What it reads | Cost | Shared |
|---|---|---|---|
catalog | Your own products in the project | Free | Never |
amazon | Amazon's catalogue through your connected seller account | Free | Never |
open_data | The Open Food Facts database, by exact barcode | Free | Yes |
page | The product page you point at | Free | Yes, when the address has no query string |
web_search | Pages on the web that state the identifier | Credits | Yes |
By default every available source is used, and a web search runs only when the free sources leave the product thin. Name the sources yourself with sources, or leave some out with excludeSources:
{ "identifiers": { "gtin": "3017620422003" }, "excludeSources": ["web_search"] }To see which sources your key can use:
curl https://ace.authoritas.com/api/v1/enrichment/lookup/sources \
-H "Authorization: Bearer $ACE_API_KEY"How web results are checked
A search can be wrong, so nothing it reports is taken on trust.
- A fact is kept only if the page it cites is one the search really visited.
- The cited pages are then read directly, and their own structured data is compared with your identifier.
- A page that states a different barcode is dropped, however alike the product looks.
Ingredients, allergens, nutrition and safety statements are never taken from a web search or from a page. They come only from a record keyed on the barcode itself.
Amazon pages are never read. An ASIN is looked up through your own connected Amazon store, and the result stays with your project.
Free lookups
GET looks up one product without charge. It reads the free sources and anything an earlier lookup already found on the web. It never runs a search.
curl "https://ace.authoritas.com/api/v1/enrichment/lookup?gtin=3017620422003" \
-H "Authorization: Bearer $ACE_API_KEY"Batches
Send up to 500 products in items. Give each one a ref and it comes back on that item's result.
{
"market": "GB",
"items": [
{ "ref": "row-1", "identifiers": { "gtin": "4006381333931" } },
{ "ref": "row-2", "identifiers": { "mpn": "WH-1000XM4" }, "hints": { "brand": "Sony" } }
]
}One bad item never fails the batch. An item that cannot be looked up comes back with "status": "error" and a code, and the response status is 207.
A small batch answers inline. A larger one returns 202 with a job, which you poll or receive by webhook:
curl https://ace.authoritas.com/api/v1/jobs/JOB_ID/results \
-H "Authorization: Bearer $ACE_API_KEY"Send an Idempotency-Key header so a retried request returns the same job. At most three lookup jobs can be queued or running at once.
Write content from what was found
Add enrich to also write content, in the same request. The copy is written only from the facts the lookup found and anything you send in product. Nothing is drawn from outside those.
{
"q": "4006381333931",
"locale": "en-GB",
"enrich": { "contentTypes": ["product-description", "meta-tags"] }
}Each item gains two fields:
contentholds the generated payload for each content type, in the same shapePOST /enrichment/contentreturns.generationsays what happened:generated,partial,failed, orskippedwith areason.
Content is written only for a product that was matched, and matched surely enough (0.8 by default, set enrich.minConfidence to change it). For anything else the status is skipped, nothing is written, and nothing is charged for writing. The reasons are MATCH_NOT_FOUND, MATCH_AMBIGUOUS, BELOW_MIN_CONFIDENCE, and NO_TITLE when a match has no product name to write about.
enrich takes the same options as the content endpoint: rules, brandVoice, config and targetLanguage. Two content types work on your own catalogue and are not written from a lookup: internal-links and custom-fields. If you ask for them they are listed under generation.skippedContentTypes and not charged.
To write about your own product with looked-up facts added, send what you hold in product. Your fields win over what the lookup found, and it is the only way a price, a stock level or a rating can reach the copy, because those are never taken from a source.
{
"identifiers": { "gtin": "4006381333931" },
"product": { "title": "STABILO point 88 fineliner, azure", "price": 1.2, "currency": "GBP" },
"enrich": { "contentTypes": ["product-description"] }
}Look up while you enrich
If you already send products to POST /enrichment/content or POST /enrichment/pipeline, add lookup and each product is looked up before its content is written. You do not call the lookup endpoint yourself.
{
"source": {
"type": "inline",
"products": [
{ "id": "sku-1", "title": "Fineliner pen, azure", "gtin": "4006381333931" }
]
},
"contentTypes": ["product-description"],
"lookup": true
}The identifiers are read from the product itself: gtin (or ean, upc, barcode), isbn, asin, mpn together with brand, and url (or link). A variant's barcode is used when the product has none of its own.
What the lookup finds is added to what you sent. It never replaces it.
- Your own product data wins wherever the two disagree.
- Only a product that was matched surely enough (
0.8by default) is written with looked-up facts. Any other product is written exactly as it would be withoutlookup. - A lookup that fails never fails the product.
lookup: true uses the free sources: your connected Amazon store, the Open Food Facts database, and the product's own page. A web search is charged, so it runs only when you name it:
{ "lookup": { "sources": ["open_data", "web_search"], "market": "GB" } }| Option | What it does |
|---|---|
sources | Use only these: amazon, open_data, page, web_search. Your own catalogue is not a source here, because the product you sent already is that record |
minConfidence | Use looked-up facts only for a match at least this sure. Default 0.8 |
freshness | maxAgeDays ignores older cached answers. refresh looks up again |
market, locale | Where the products are sold, and the language to look them up in |
Each product's result gains a lookup field that says what happened:
{
"productId": "sku-1",
"content": { "product-description": { /* … */ } },
"lookup": {
"status": "grounded",
"confidence": 0.95,
"factCount": 6,
"sources": [{ "id": "open_data", "status": "miss" }, { "id": "web_search", "status": "hit" }]
}
}| Status | Meaning |
|---|---|
grounded | Facts were found and the content was written with them |
not_matched | The product was looked up and nothing sure enough came back |
no_identifier | Nothing on the product could be looked up |
failed | The lookup itself failed. The content was still written |
The response also carries lookup totals and any attributions the sources ask for. For a job, both are in meta on GET /jobs/{id}/results. A request that looks products up counts as more work, so it becomes a job a little sooner than the same request without lookup.
In the dashboard
Feed enrichment has the same option. In the enrichment setup, under the AI options, turn on Look products up. Each product in the run is looked up by its barcode, part number or page before anything is generated.
- An attribute a published source states (colour, material, pattern, brand, gender, age group) is filled in where yours is empty. A value you already have is never replaced, even when the run is allowed to overwrite.
- Everything else the lookup finds, such as a description or a pack size, is given to the writers as a fact. It is never copied into your feed.
- In Review, each value that was looked up shows its Source under it, with a link to the page that states it. You decide whether to publish it, field by field, as with any other change.
Look products up on its own uses free sources. Also search the web finds more and uses credits for each product searched. A product already looked up is served from cache and not charged again.
Scheduled and automatic runs look products up with the free sources only. They never search the web, because nobody is watching what they spend. A run you start yourself, including Run now, searches as your setup says.
A setup saved with the older "MPN / GTIN lookup" option ticked keeps it: that option now does what it said.
Check before you run
dryRun validates the request, says which sources would run for each item, and estimates the cost. Nothing is looked up and nothing is charged.
{ "items": [{ "q": "4006381333931" }, { "q": "not-a-barcode-1234" }], "dryRun": true }Cost
A web search is charged, and so is writing content with enrich. Everything else is free.
- A lookup that needs a web search typically costs about 16 credits.
- Content written with
enrichcosts what it costs on the content endpoint: one unit per product per content type that is written. lookupon the content and pipeline endpoints adds nothing with the free sources. Withweb_searchnamed, each search is charged as it is here.- An answer that was already found, by you or by anyone else, is served from cache and costs nothing. Each item's
usage.cacheHitsand theX-Lookup-Cacheresponse header tell you when that happened. - If a search runs and we then fail to return its result, you are not charged for it.
Test keys can use every source. Free sources use none of the monthly quota. An item that may run a web search uses 5 units, and each content type written for an item uses 1. The same holds for lookup on the content and pipeline endpoints: 5 units for each product that may be searched for, on top of the units for its content.
Showing the data
Some sources ask for credit wherever their data is shown. The response lists what is owed, once per source, under attributions, and each fact carries its own licence and attribution.
- Open Food Facts data is published under the Open Database License. Its images are under Creative Commons Attribution ShareAlike. If you republish either, show the attribution text.
- Web results come with the address of the page that states each fact. If you show the fact to your own users, link to that page.
Errors
| Code | Status | Meaning |
|---|---|---|
NO_IDENTIFIER | 400 or per item | Nothing was sent that can be looked up |
INVALID_IDENTIFIER | 400 or per item | What was sent failed validation, for example a wrong check digit. details says why |
SOURCE_UNAVAILABLE | 400 | None of the sources you named is available to this key |
LOOKUP_TIMEOUT | per item | The request ran out of time before this item. Send fewer items, or use "mode": "async" |
LOOKUP_FAILED | per item | The lookup could not be completed. It is safe to retry |
LOOKUP_CONCURRENCY_LIMIT | 429 | Three lookup jobs are already queued or running |
An identifier that could not be used as sent, while another one could, is not an error. It is listed under that item's warnings.
Limits
- 500 items per request. Up to 10 with
"mode": "sync". - The Open Food Facts source is shared by all callers and rate limited by that service. Under load it reports
unavailablewithRATE_LIMITEDfor some items, andmatch.partialis set. Ask again later and the answer is served from cache. - A part number is searched on the web only when you send its brand.
- A stock code and a store-internal barcode are matched in your own catalogue and nowhere else.