Skip to content
Iknow

80,000 Recipes, One Better Taxonomy: Sharpening Search Across a Global Manufacturing Knowledge Base.

Global pharmaceutical company focused on prescription medicines and vaccines · Pharmaceuticals & Biotechnology

Refining a Pharmaceutical Manufacturing Taxonomy to Distinguish Large- and Small-Molecule Drug Modalities

Executive summary

Company Z’s manufacturing division formulates, packages, and distributes products to more than 140 markets worldwide. In prior engagements, Iknow had built Company Z’s primary enterprise taxonomy from the ground up: more than 1,840 terms across a four-level hierarchy and 10 facets, complete with synonyms, acronyms, and alternative labels for drugs at different development and commercialization stages, and deployed with autoclassification rules across all of Company Z’s SharePoint repositories. As Company Z’s product mix, like the broader pharmaceutical industry, shifted increasingly toward complex biologics alongside traditional small-molecule drugs, Company Z’s KM Center of Excellence wanted the taxonomy’s manufacturing and commercialization facets to draw a sharper line between large-molecule-specific, small-molecule-specific, and generic process steps. Company Z asked Iknow for continued support to refine the taxonomy and improve autoclassification accuracy by drug modality.

Iknow reviewed user feedback submitted through Company Z’s intranet comments channel, analyzed tagging patterns across 10 large content collections and roughly 80,000 entries in Company Z’s manufacturing “recipe” database, and conducted stakeholder interviews to identify where terms were over- or under-applied. Iknow used those findings to add about 200 new taxonomy terms, reorganize the manufacturing and commercialization facets to provide clearer separation by molecule type, and implement and test the revised model in Semaphore Ontology Editor before publishing it. The work was completed on schedule, with testing confirming significant improvements in tagging accuracy for detailed manufacturing terms.

Background & context

About the Client

Company Z is a global pharmaceutical manufacturer. It sells products in more than 140 countries. Company Z’s manufacturing division oversees the formulation, packaging, and distribution of its global product portfolio. Its Knowledge Management Center of Excellence has worked with Iknow on multiple engagements to build and continuously refine the enterprise taxonomy that powers search across Company Z’s SharePoint repositories, helping Company Z’s scientists and engineers find relevant content quickly and precisely, regardless of the technical discipline from which it originates.

Industry Context

Pharmaceutical manufacturing is in the midst of a structural shift toward biologics — large, complex molecules such as monoclonal antibodies produced in living cells — alongside traditional small-molecule drugs made through chemical synthesis, which still account for most approved medicines. The global biologics manufacturing market was valued at roughly $40 billion in 2025 and is projected to grow nearly 17% annually through 2033, driven in part by biologics’ far greater manufacturing complexity. While a small molecule is produced through relatively straightforward chemical synthesis, a biologic requires cell line development, bioreactor cultivation, and specialized downstream purification. A taxonomy that lumps both together under generic process terms becomes less useful precisely as this mix shifts, since a scientist searching for guidance on a bioreactor harvest step needs very different content than one working on a small-molecule tablet formulation.

Current Situation

Company Z’s existing taxonomy — more than 1,840 terms across a four-level hierarchy and 10 facets, already refined once based on testing, end-user feedback, and analysis of content such as manufacturing deviation reports — still did not clearly distinguish large-molecule-specific, small-molecule-specific, and generic steps within its manufacturing and commercialization facets. Company Z asked Iknow to further improve metadata tagging quality by drug modality, review user feedback and taxonomy term quality more broadly, review classification results for additional improvement opportunities, and implement and test the required changes.

Problem / challenge

  • An already-mature taxonomy still blends distinct drug modalities. Even Company Z’s robust 1,840-term taxonomy didn’t clearly separate large-molecule-specific, small-molecule-specific, and generic manufacturing and commercialization steps, limiting how precisely users could filter by drug modality.
  • A widening gap with Company Z’s evolving product mix. As Company Z’s portfolio, like the broader industry, incorporated more large-molecule biologics alongside traditional small-molecule drugs, the taxonomy’s ability to distinguish between the two became more important than when it was first built.
  • Feedback was scattered across an informal comments channel. A clear signal about inaccurate tagging appeared in Company Z’s intranet comments, but only as unstructured commentary that needed to be synthesized into concrete taxonomy changes.
  • A massive, unstructured evidence base to mine. Confirming and expanding modality-specific terms required analyzing tagging patterns across 10 large content collections and roughly 80,000 entries in Company Z’s manufacturing recipe database — far too much to review manually.

Project objectives

  • Improve metadata tagging quality for large and small molecules, particularly in the taxonomy’s manufacturing and commercialization facets, enabling more precise searches by drug modality.
  • Review user feedback and refine the taxonomy’s term definitions, synonyms, acronyms, and other alternative labels.
  • Review classification results and identify additional opportunities to improve autoclassification accuracy.
  • Implement and test the required changes to the taxonomy model and classification rules.

Iknow’s approach

How Iknow Structured the Work

Iknow structured the engagement into four phases, moving from user feedback and data analysis to stakeholder validation, then to staged implementation and testing of the revised taxonomy.

Key Activities & Decisions

  • Review of user feedback. Iknow reviewed user feedback on the taxonomy submitted via Company Z’s intranet comments channel, focusing on comments about the accuracy and balance of tagging between small and large molecules.
  • Classification review and data analysis. Iknow compiled current tagging reports for 10 large content collections, showing tag frequency in the Manufacturing Step and Commercialization Step facets by molecule type, to identify over- and under-tagging of specific terms. Iknow also analyzed Company Z’s manufacturing “recipes” database of roughly 80,000 entries to identify small- and large-molecule-specific steps and potential new taxonomy terms.
  • Stakeholder interviews. Iknow conducted user interviews to explore potential taxonomy improvements, prompting adjustments to the taxonomy’s hierarchy and topic grouping, and surfacing many new terms for greater search granularity.
  • Implementation and testing. Iknow first edited the taxonomy model in Excel for quick user review, then implemented the changes in Semaphore Ontology Editor, and performed classification testing on the revised model before final publication.

Stakeholders & Collaboration

Iknow served as prime contractor, working directly with Company Z’s KM Center of Excellence and drawing on both intranet user feedback and direct stakeholder interviews with scientists and engineers, thereby continuing Iknow’s established multi-engagement taxonomy relationship with Company Z.

Challenges & how Iknow overcame them

Finding Real Signal in 80,000 Unstructured Entries

Confirming genuine large- and small-molecule-specific terms required more than intuition — it required evidence from Company Z’s operational data. Iknow addressed this by systematically analyzing the manufacturing recipe database alongside tagging-frequency reports from 10 large content collections, cross-referencing to identify where existing tags were over- or under-applied by molecule type, and then proposing new terms, grounding every change in actual usage data rather than guesswork.

Turning Informal Feedback into a Validated Taxonomy Revision

Scattered intranet comments alone could not tell Iknow which changes would genuinely improve search without disrupting an already-deployed, actively used taxonomy. Iknow addressed this by combining structured review of user feedback with direct stakeholder interviews, then staging implementation — an Excel draft for quick user review, followed by Semaphore Ontology Editor for implementation and classification testing — before publishing the final model.

Results & impact

Quantitative Outcomes

  • New terms added: approximately 200 new taxonomy terms, along with corresponding synonyms and acronyms.
  • Data analyzed: tagging reports from 10 large content collections and a manufacturing recipe database with roughly 80,000 entries.
  • Delivery: completed on schedule within the four-month contract period.

Qualitative Outcomes

Several changes to the taxonomy’s groupings and hierarchy improved the separation of small-molecule, large-molecule, and generic process steps, and testing showed significant gains in tagging accuracy for detailed manufacturing terms. The new terms and reorganized facets gave Company Z’s scientists and engineers a way to filter search results by drug modality that hadn’t existed before, reducing noise from irrelevant results when someone working on a biologic searched for a manufacturing step that also applied, under a different process, to small-molecule production. The engagement also extended Iknow’s multi-year, iterative taxonomy partnership with Company Z, keeping the underlying search infrastructure aligned with Company Z’s product mix as it shifts further toward complex biologics — a trend playing out across the broader pharmaceutical manufacturing industry rather than one unique to Company Z.

Timeline to Impact

Iknow delivered all four phases — feedback review, classification analysis, stakeholder interviews, and implementation and testing — within the four-month contract period.

Iknow’s capabilities demonstrated

Core Skills

  • Taxonomy refinement & governance
  • Classification analytics
  • Autoclassification rule tuning
  • Stakeholder-driven requirements gathering

Methods & Frameworks

  • Tagging-frequency analysis by content collection
  • Structured user feedback review
  • Stakeholder interviews
  • Staged model implementation and testing

Technologies & Tools

  • Progress Semaphore Ontology Editor
  • SharePoint enterprise search and manufacturing recipe database analysis

Put this experience to work on your problem.

Much of our work never reaches the website. Book a call, tell us your sector and we will walk you through the engagements that map to yours.