database-lookup

Query documented public database APIs with explicit endpoints, filters, pagination, and provenance. Use when a scientific, regulatory, financial, or other database-backed fact must be retrieved reproducibly from a named source rather than inferred from general knowledge.

Install

Hot:34

Download and extract to your skills directory

Copy command and send to AI Agent for auto-install:

Download and install this skill https://openskills.cc/api/download?slug=k-dense-ai-skills-database-lookup&locale=en&source=copy

Database Lookup – Public Database API Query Tool

Skill Overview


Database Lookup is an intelligent query tool cataloging 78 public databases. Through documented API access patterns, it helps you reproducibly retrieve factual data from authoritative sources in fields such as science, regulation, and finance.

Use Cases

1. Scientific Research and Data Analysis


When you need to obtain research data from authoritative databases, Database Lookup can help you with:

  • Biological research: Query gene information (NCBI Gene, Ensembl), protein sequences (UniProt), metabolic pathways (Reactome, KEGG), gene expression data (GEO, GTEx), and disease variants (ClinVar, COSMIC)

  • Chemistry and pharmaceuticals: Obtain compound properties (PubChem), bioactivity data (ChEMBL), drug labels (FDA, DailyMed), and molecular similarity searches (ZINC)

  • Physics and astronomy: Retrieve near-Earth asteroid data (NASA NeoWS), exoplanet data, atomic spectra (NIST), and astronomical observations (SDSS, SIMBAD)

  • Materials science: Query material band structures and crystal structures (Materials Project, COD)
  • 2. Financial, Economic, and Compliance Data


    When your business requires accurate official data:

  • Economic indicators: U.S. economic time series (FRED), employment data (BLS), GDP data (BEA), and World Bank development indicators

  • Financial regulation: SEC company filings, U.S. patent searches (USPTO), and Federal Reserve data

  • Demographics: U.S. Census data, Eurostat statistics, and WHO health indicators
  • 3. AI Agent and Automation Integration


    Provide AI Agents with reliable data access capabilities:

  • Identifier resolution: Automatically convert gene symbols, compound names, disease names, and other identifiers between database-specific formats

  • Completeness validation: Support paginated queries and count verification to ensure retrieval completeness

  • Provenance auditing: Return provenance information such as database endpoints, parameters, and access dates, making results reproducible

  • Security safeguards: Handle rate limits, API key management, and input validation to protect your API access
  • Core Features

    1. Authoritative Database Selection and Querying


    Covers 78 public databases and automatically selects the most authoritative data sources for you:

  • Multidomain coverage: Biology, chemistry, physics, earth sciences, materials science, economics and finance, social sciences, patents, and regulation

  • Intelligent selection guidance: Recommend primary and supplementary databases based on your query intent, avoiding unnecessary API calls

  • Cross-validation: Support cross-validation across multiple databases for critical queries to ensure data accuracy
  • 2. Identifier Conversion and Resolution


    Automatically handle differences in identifier formats across databases:

  • Gene identifiers: Support conversion among gene symbols, NCBI Gene IDs, Ensembl IDs, and UniProt IDs

  • Compound identifiers: Support compound names, CAS numbers, PubChem CIDs, ChEMBL IDs, SMILES, and InChIKeys

  • Variant data: Support rsIDs, genomic coordinates, and clinical variant identifiers

  • Disease identifiers: Support disease names, MONDO IDs, EFO IDs, and HPO terms

  • When a query fails, automatically try alternative identifiers or databases
  • 3. Completeness Validation and Provenance Auditing


    Ensure the reliability and reproducibility of data retrieval:

  • Count verification: For exhaustive retrievals, count records first, estimate the cost, and paginate until all data has been retrieved

  • Pagination handling: Support multiple pagination modes, including offset/limit, cursors, and page numbers

  • Provenance records: Return complete provenance information, including database endpoints, query parameters, access dates, identifier conversions, and count reconciliation

  • Warnings: Clearly flag incomplete pagination, ambiguous filters, outdated data, source limitations, and other issues

  • Handling untrusted data: Treat API responses as third-party data and provide recommendations for secure handling
  • Frequently Asked Questions

    Which databases does Database Lookup support?

    It currently supports 78 public databases across 9 major domains:

    Biology and genomics (31): Reactome, KEGG, UniProt, STRING, Ensembl, NCBI Gene/Protein/Taxonomy, GEO, GTEx, PDB, AlphaFold, EMDB, InterPro, BioGRID, Gene Ontology, QuickGO, dbSNP, SRA, gnomAD, UCSC Genome, ENCODE, JASPAR, Human Protein Atlas, Human Cell Atlas, LINCS L1000, RummaGEO, PRIDE, Metabolomics Workbench, BRENDA, MouseMine, ENA, Addgene, ChEBI

    Chemistry and pharmaceuticals (9): PubChem, ChEMBL, DrugBank, FDA, DailyMed, KEGG, ZINC, BindingDB, ChEBI

    Physics and astronomy (5): NASA NeoWS, NASA Exoplanet Archive, NIST, SDSS, SIMBAD

    Earth and environmental sciences (4): USGS, NOAA, EPA, OpenWeatherMap

    Diseases and clinical research (12): Open Targets, COSMIC, ClinPGx, ClinicalTrials.gov, OMIM, ClinVar, GDC/TCGA, cBioPortal, DisGeNET, GWAS Catalog, Monarch Initiative, HPO

    Patents and regulation (2): USPTO, SEC EDGAR

    Economics and finance (9): FRED, Federal Reserve, BEA, BLS, World Bank, ECB, U.S. Treasury, Alpha Vantage, Data Commons

    Materials science (2): Materials Project, COD

    Social sciences (3): U.S. Census, Eurostat, WHO GHO

    How do I configure API keys?

    Some databases require API keys for access or higher rate limits. Database Lookup supports configuration through the following environment variables:

    API keys available through free registration:

  • FRED_API_KEY: https://fred.stlouisfed.org/docs/api/api_key.html

  • BEA_API_KEY: https://apps.bea.gov/API/signup/

  • BLS_API_KEY: https://data.bls.gov/registrationEngine/

  • NCBI_API_KEY: https://www.ncbi.nlm.nih.gov/account/settings/

  • OPENFDA_API_KEY: https://open.fda.gov/apis/authentication/

  • PATENTSVIEW_API_KEY: https://patentsview.org/apis/keyrequest

  • NASA_API_KEY: https://api.nasa.gov (the free DEMO_KEY can be used)

  • NOAA_API_KEY: https://www.ncdc.noaa.gov/cdo-web/token

  • OPENWEATHERMAP_API_KEY: https://openweathermap.org/appid

  • MP_API_KEY: https://materialsproject.org (free account)
  • No key required or restricted access:

  • DrugBank and COSMIC require paid access or a license

  • Most databases support access without a key, but with lower rate limits
  • To configure them, add the keys to a .env file or to your environment variables. Database Lookup will detect and use them automatically.

    How are identifier format mismatches handled?

    Different databases use different identifier systems. When a query fails, Database Lookup automatically tries the following strategies:

    Examples of identifier conversion:

  • Gene symbol "TP53" → NCBI Gene ID "7157" → Ensembl ID "ENSG00000141510" → UniProt "P04637"

  • Compound name "aspirin" → PubChem CID "2244" → ChEMBL ID "CHEMBL25" → ZINC ID "ZINC000000000053"

  • Variant "rs334" can be used directly in dbSNP, ClinVar, GWAS Catalog, and gnomAD
  • Automatic conversion workflow:

  • First query the primary database using the identifier you provided

  • If that fails, look it up in a reference database, such as NCBI Gene

  • After obtaining a standard ID, convert it to the identifier format used by the target database

  • If it still fails, try an alternative database
  • You can explicitly specify the identifier type, such as “query using an Ensembl ID,” to avoid ambiguity.

    How can I ensure retrieval completeness?

    For exhaustive retrievals, such as “all clinical trials” or “all known variants,” Database Lookup takes the following measures:

    Completeness guarantees:

  • Count first: Use count endpoints or metadata to obtain the total number of records

  • Estimate cost: Estimate the number of API calls based on the record count and page size

  • Confirm boundaries: Request confirmation when there are more than 10,000 records or 100 API calls

  • Establish ordering: Use sorting or stable cursors to ensure consistent ordering

  • Record batches: Record the number requested and returned for each page or cursor

  • Reconcile counts: Compare the expected total, server-reported total, and total remaining after local filtering

  • Make failures visible: Clearly report limitations if pagination stops prematurely or counts do not match
  • Provenance information:

    Each query returns:

  • The database and endpoint queried

  • Request parameters and filters

  • Identifier conversion records

  • Count reconciliation (expected/retrieved/remaining after filtering)

  • Filters applied locally

  • Completeness warnings
  • You can use this information to reproduce the entire retrieval process.

    What are the API rate limits?

    Rate limits vary by database. Database Lookup handles rate limits automatically.

    NCBI services (Gene, GEO, Protein, Taxonomy, dbSNP, SRA):

  • Without a key: 3 requests/second

  • With a key: 10 requests/second
  • Other common limits:

  • Ensembl: 15 requests/second

  • SEC EDGAR: 10 requests/second

  • NOAA: 5 requests/second (with a token)

  • BLS v1: 25 requests per day (without a key)
  • Automatic handling:

  • Serialize requests to APIs with rate limits

  • Wait and retry when HTTP 429/503 responses are encountered

  • Recommend using API keys to increase quotas
  • For bulk retrievals, Database Lookup prioritizes official bulk downloads or database dumps.

    Who is Database Lookup suitable for?

    Researchers: Obtain authoritative scientific data for papers, experimental design, and data validation

    Bioinformatics researchers: Query multidimensional data on genes, proteins, metabolic pathways, disease variants, and more

    Drug development professionals: Retrieve compound properties, bioactivity, drug labels, and target information

    Financial analysts: Obtain economic indicators, SEC filings, interest rate data, and other financial regulatory data

    AI Agent developers: Integrate reliable data sources into Agents with identifier resolution, completeness validation, and provenance auditing

    Data engineers: Quickly integrate multiple public database APIs and handle complex query and pagination logic

    Compliance professionals: Query regulatory data on patents, trademarks, FDA labels, clinical trials, and more

    Whether you need a one-time query or a bulk integration, Database Lookup can provide reproducible and auditable data retrieval results.