All case studies

SEC Patent Licensing Extraction

A multi-source research system that linked government-funded patents, award records, public licensing evidence, and company data into a reviewable analysis layer.

Client
Health Economics Research Consultancy
Scope
Patent, funding, licensing, and company-data integration
Outcome
Reusable evidence base for an economic-impact study

Research brief

The consultancy needed a reproducible way to connect patents containing federal-funding disclosures with the awards that supported them, public evidence of subsequent licensing, and company-level information. The resulting research system supported a broader economic and policy study while keeping every derived relationship connected to its source evidence.

Patent and funding identification

Clarivate Derwent Patent Search provided the initial patent set through its Government Interest field. Government Interest Statements, USPTO patent information, and Certificates of Correction were then reviewed to isolate patents with relevant U.S. government-funding disclosures. The rules excluded foreign-government support, statements that no government funding applied, and records where the government-interest section was not applicable. Statements describing federal support or the U.S. government’s royalty-free license for governmental purposes were retained.

A model accessed through the OpenAI API extracted federal agencies and award identifiers from the retained statements and produced normalized patent-funding pairs. NIH-funded links were supplemented with the NIH RePORTER patents file. Relevant award records were then retrieved from the NIH RePORTER API and USAspending Award Data Archive files, allocated across each award’s active years, and joined back to the linked patents.

Licensing evidence and company matching

The SEC EDGAR API was searched for licensing-related filings that could reveal agreements unavailable through the patent source. After deterministic search and filtering narrowed the documents, a model accessed through the OpenAI API reviewed the full text to identify agreements involving patents connected to federal funding. Supporting filing passages remained attached to the extracted records for verification.

Patent assignee names from Clarivate and licensee names from the SEC extraction were uploaded through PitchBook’s list-matching workflow to obtain firm-level employment and financial information. When a name did not resolve directly, company websites were used to locate and manually verify the appropriate PitchBook record. Alternate names, subsidiaries, and uncertain matches were handled as explicit review cases instead of being joined automatically.

Reliability controls

Automation generated candidate relationships; it did not convert every textual match into a verified record. Inclusion and exclusion rules were applied before model extraction, source passages were retained alongside structured fields, and exception queues separated ambiguous award identifiers, company matches, and licensing language for manual review.

The pipeline also recorded source, study year, processing status, and match method. Separate adapters for Clarivate, USPTO information, NIH RePORTER, USAspending, SEC EDGAR, and PitchBook made it possible to update one source or rule without rebuilding the full research system.

Outcome

The engagement produced a reusable evidence base for an economic-impact study. Researchers could work from a consistent set of patent, funding, licensing, and company relationships while returning to the underlying statement, award record, filing passage, or matching decision whenever a result required review.

The modular architecture also left room for later study updates, additional fields, revised matching rules, and new output tables without discarding the original provenance record.

Have a similar workflow or research question?

Discuss a project