{"id":2454,"date":"2026-08-08T01:46:10","date_gmt":"2026-08-07T22:46:10","guid":{"rendered":"https:\/\/picajet.com\/articles\/glossary\/full-text-search\/"},"modified":"2026-08-08T03:46:06","modified_gmt":"2026-08-08T00:46:06","slug":"full-text-search","status":"publish","type":"glossary","link":"https:\/\/picajet.com\/articles\/glossary\/full-text-search\/","title":{"rendered":"Full-text search"},"content":{"rendered":"<p class=\"wp-block-paragraph\">Full-text search treats every indexed field as fair game for a keyword match, rather than restricting a query to a single tag or title field. That difference matters in a DAM because assets accumulate text in places nobody tagged on purpose \u2014 a filename someone typed years ago, a caption pasted from a wire service, a description field an editor half-filled in.<\/p><p class=\"wp-block-paragraph\">The practical value shows up with legacy content. A library migrated from an old file server rarely has consistent keyword tags, but it usually has filenames and folder paths carrying real information. Full-text search picks that up without anyone re-tagging a single asset, which is often the only reason a large historical archive is searchable at all on day one.<\/p><p class=\"wp-block-paragraph\">What it does not do automatically is read the pixels of an image. Text visible on a product label, a presentation slide exported as an image, or a scanned contract only becomes searchable once OCR has extracted it and that extracted text has been fed into the same index full-text search queries against \u2014 two separate steps that are easy to assume happened when only one did.<\/p>","protected":false},"excerpt":{"rendered":"<p>Search that scans and matches words across every indexed field \u2014 filenames, captions, embedded metadata, OCR output \u2014 rather than a single tag field, returning anything containing the query term.<\/p>\n","protected":false},"author":0,"featured_media":0,"template":"","meta":{"footnotes":"","faq":[{"question":"What does full-text search index that a single keyword-tag search doesn't?","answer":"Every indexed field \u2014 filenames, captions, embedded metadata, and OCR output \u2014 rather than just one tag field, so it returns anything containing the query term wherever that text happens to live on the asset."},{"question":"Why is full-text search especially valuable for a legacy library migrated from an old file server?","answer":"Such a library rarely has consistent keyword tags, but it usually has filenames and folder paths carrying real information, which full-text search picks up without anyone re-tagging a single asset."},{"question":"Does full-text search automatically make text inside scanned images or photos searchable?","answer":"No \u2014 text visible on a product label, a slide exported as an image, or a scanned contract only becomes searchable once OCR has extracted it and that extracted text has been fed into the same index, which is a separate step that's easy to assume already happened."},{"question":"What fields does full-text search typically cover by default?","answer":"Filenames and the caption\/description field are indexed by default; custom metadata fields are indexed if configured; the OCR text layer is indexed only if OCR has been run and connected to the index."},{"question":"Why does full-text search compensate for gaps in manual tagging?","answer":"It draws on whatever raw text already sits on or around the asset \u2014 filename, caption, embedded metadata \u2014 instead of requiring someone to have anticipated the exact search term as a keyword tag."},{"question":"What's a common mistake teams make about OCR and full-text search?","answer":"Assuming full-text search automatically covers OCR text from images and scanned PDFs, when in most DAMs OCR has to be run and indexed as a separate step \u2014 so a scanned document with clearly visible text can return zero results until that pipeline is switched on."}],"checked_date":"2026-08-07","sources":[],"kicker":"","fact_checker":0,"reading_time":0,"revisions":[],"seo_title":"","seo_description":"","noindex":false,"related":[2467,2458,2465,2569,2450,2402],"definition":"Search that scans and matches words across every indexed field \u2014 filenames, captions, embedded metadata, OCR output \u2014 rather than a single tag field, returning anything containing the query term.","why":"Full-text search is what makes an old scanned brochure findable by a phrase visible in its own text, or a photo findable by a word buried in its caption rather than its keyword tags. It compensates for gaps in manual tagging by drawing on whatever raw text already sits on or around the asset, instead of requiring someone to have anticipated the exact search term as a keyword.","example_rows":[{"field":"Filename","values":"indexed by default"},{"field":"Caption \/ description field","values":"indexed by default"},{"field":"Custom metadata field","values":"indexed if configured"},{"field":"OCR text layer","values":"indexed only if OCR is run and connected to the index"}],"mistake":"Teams assume full-text search automatically covers OCR text pulled from images and scanned PDFs, when in most DAMs OCR has to be run and indexed as a separate step \u2014 so a scanned document with clearly visible text can still return zero results until that pipeline is switched on.","deep_link":""},"silo":[24],"class_list":["post-2454","glossary","type-glossary","status-publish","hentry","silo-glossary"],"_links":{"self":[{"href":"https:\/\/picajet.com\/articles\/wp-json\/wp\/v2\/glossary\/2454","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/picajet.com\/articles\/wp-json\/wp\/v2\/glossary"}],"about":[{"href":"https:\/\/picajet.com\/articles\/wp-json\/wp\/v2\/types\/glossary"}],"version-history":[{"count":3,"href":"https:\/\/picajet.com\/articles\/wp-json\/wp\/v2\/glossary\/2454\/revisions"}],"predecessor-version":[{"id":3574,"href":"https:\/\/picajet.com\/articles\/wp-json\/wp\/v2\/glossary\/2454\/revisions\/3574"}],"wp:attachment":[{"href":"https:\/\/picajet.com\/articles\/wp-json\/wp\/v2\/media?parent=2454"}],"wp:term":[{"taxonomy":"silo","embeddable":true,"href":"https:\/\/picajet.com\/articles\/wp-json\/wp\/v2\/silo?post=2454"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}