Skip to content
Iknow

Applying Natural Language Processing and Custom Rule Sets to Turn Unstructured Job Data into Trend Reporting.

Independent technology research and advisory firm · Software & Services

Mining 8,000 Job Postings to Reveal Where the BI Talent Market Was Heading

Executive summary

Company M wanted to analyze job postings in the business intelligence (BI) field over two years to report on relevant trends, but the information it needed was buried in unstructured job posting text, requiring sophisticated natural language processing tools to extract. Company M hired Iknow to mine, cleanse, and analyze the data. Indeed provided more than 8,000 job postings for the project, each containing at least one of eight BI-related phrases in the job title, with roughly half collected in June 2007 and half in June 2008. Iknow’s scope was to extract geographic locations, industry sectors, BI vendor names, salary ranges, skill levels, and time trends from the unstructured text, cleanse and standardize the results, and present the information in a set of charts.

Over roughly three months, Iknow executed a four-step pipeline: preprocessing to clean and structure the raw job posting data, entity extraction using SAP BusinessObjects Text Analysis software with custom name catalogs and rule sets, post-processing in a Microsoft SQL Server database to standardize and cut the data, and Xcelsius chart creation to present results across eight dimensions, including geography, industry, top hiring companies, revenue, salary, BI vendors, job categories, and skills. Iknow delivered two full source-data spreadsheets, 29 summary spreadsheets, and nine summary PowerPoint files of charts — the complete set of outputs Company M needed to prepare its trend report.

Background & context

About the Client

Company M is one of the world’s most prominent independent technology and business research and advisory firms, known for its vendor-neutral analysis.

Industry Context

By 2007 and 2008, business intelligence had become one of the fastest-growing categories within enterprise IT, driving rising demand for specialized BI roles that traditional labor-market data sources were often slow to capture. Online job boards had begun accumulating large volumes of real-time hiring data offering a far more current window into an emerging labor market than survey-based research, but that data lived entirely in unstructured free text, invisible to conventional quantitative analysis without natural language processing. For a research firm like Company M, known for closely tracking enterprise technology adoption, text mining actual job postings offered a more evidence-grounded way to report on where the BI talent market was heading, rather than relying on surveys or analyst judgment alone.

Current Situation

Company M wanted to analyze job postings in the BI field over two years so that it could report on relevant trends, but because the desired information was buried inside the text of the job postings, sophisticated natural language processing tools were necessary to extract the data. Indeed provided more than 8,000 job postings containing any of eight BI-related phrases in the job title, split roughly evenly between June 2007 and June 2008, and Company M hired Iknow to mine, cleanse, and analyze the data.

Problem / challenge

  • Desired data locked inside unstructured text. The BI labor-market data Company M needed existed only as unstructured free text buried inside more than 8,000 job postings, invisible to conventional quantitative analysis.
  • Inconsistent, incomplete, and irrelevant source data. Job postings varied widely in format and completeness — inconsistent company names, missing data, and non-BI postings mixed in with genuinely relevant ones.
  • Highly variable formats for the same underlying data. The desired data types each appeared in multiple formats and units within the free text, such as revenue figures in five different currencies and salary information presented inconsistently.
  • No path from raw extraction to publishable analysis. Company M needed the extracted data cleansed, standardized, and translated into executive-ready charts, not just a raw extraction.

Project objectives

  • Extract geographic, industry, vendor, salary, skill, and time-trend data from more than 8,000 unstructured BI job postings.
  • Cleanse and standardize the extracted results into a consistent, analyzable dataset.
  • Analyze and present the resulting information in a set of clear charts spanning multiple dimensions of the BI job market.
  • Compare hiring trends between June 2007 and June 2008 to surface meaningful year-over-year change.

Iknow’s approach

How Iknow Structured the Work

Iknow structured the engagement as a four-step pipeline performed in series — preprocessing, entity extraction, post-processing, and chart creation — reflecting Iknow’s automated content classification methodology for turning a large volume of unstructured text into clean, decision-ready analysis.

Key Activities & Decisions

  • Preprocessing. Iknow converted the XML input files into Excel spreadsheets, deleted spurious non-BI job descriptions, standardized inconsistent company names, and manually corrected missing or erroneous data where determinable from the job text, also defining new Industry (NAICS code) and Job Category fields across all 8,000-plus records.
  • Entity extraction. Using SAP BusinessObjects Text Analysis software, Iknow ran an iterative extraction process, building custom Name Catalogs for BI vendors, products and tools, and skills, along with rule sets to extract company revenue (across five currencies and multiple numeric formats) and salary information, including ranges and contractor wages.
  • Terminology rule sets. Iknow also built rule sets to identify structured-data terminology, such as SQL and OLAP variants, and unstructured-data terminology, such as text mining and entity extraction.
  • Post-processing and standardization. Iknow migrated the cleansed data into a Microsoft SQL Server database, built joins and views to generate different data cuts, eliminated duplicate extractions, converted all revenue to U.S. dollars, and converted salary ranges to mean values.
  • Dashboard and chart creation. Iknow used the resulting spreadsheets to build multiple Xcelsius projects, creating dashboards displaying results by geography, industry, top hiring companies, revenue, salary, BI vendors, job categories, and skills.

Stakeholders & Collaboration

Iknow served as prime contractor, working directly with Company M to deliver the complete data extraction, cleansing, and visualization outputs needed to support Company M’s own published trend report.

Challenges & how Iknow overcame them

Normalizing Highly Inconsistent Data at a Volume of 8,000-Plus Records

Free-text job postings varied enormously in company naming, completeness, and relevance, and doing that cleanup manually at scale risked being both slow and inconsistent. Iknow addressed this by building a disciplined preprocessing step — removing spurious postings, standardizing company names, and fixing missing or erroneous data — before any automated extraction began, ensuring the entity extraction step operated on a clean, consistent foundation.

Extracting Meaningful Entities Without Relying on Generic Software Alone

Out-of-the-box text analysis software was not built to recognize BI-specific vendors, products, skills, or the many different formats used to express revenue and salary in free text. Iknow addressed this by building custom Name Catalogs and rule sets tailored specifically to the vocabulary and formats found in these job postings, refined through an iterative extraction process rather than accepting the software’s default results.

Results & impact

Quantitative Outcomes

  • Source data volume: More than 8,000 job postings analyzed, spanning June 2007 and June 2008.
  • Data normalization scope: Data standardized across five currencies and multiple revenue and salary formats.
  • Dashboard categories delivered: Xcelsius dashboards delivered across eight dimensions: geography, industry, companies, revenue, salary, BI vendors, job categories, and skills.
  • Deliverables produced: Final outputs: 2 full source-data spreadsheets, 29 summary spreadsheets, and 9 summary PowerPoint files of charts.
  • Engagement duration: Three-month assignment.

Qualitative Outcomes

Iknow analyzed the job descriptions and created the outputs that Company M needed to prepare its report, direct evidence that the engagement’s outputs were fit for Company M’s own downstream publication purposes. By combining rigorous preprocessing, custom entity extraction, relational post-processing, and executive-ready dashboards, Iknow turned a large, messy body of unstructured text into a clean, multi-dimensional, decision-ready dataset — giving Company M an evidence base for BI labor-market trend reporting grounded in real job-posting data rather than survey responses alone.

Timeline to Impact

During the three-month engagement, Iknow delivered the complete dataset, summary spreadsheets, and dashboard-ready chart outputs, providing Company M with everything it needed to begin preparing its trend report.

Iknow’s capabilities demonstrated

Core Skills

  • Text mining and natural language processing
  • Data extraction, cleansing, and standardization
  • Business intelligence and data visualization
  • Custom entity extraction rule development

Methods & Frameworks

  • Four-step structured text mining pipeline (preprocessing, extraction, post-processing, visualization)
  • Iterative rule refinement
  • Relational database-driven data cutting

Technologies & Tools

  • SAP BusinessObjects Text Analysis; Xcelsius (SAP Dashboards)
  • Microsoft SQL Server; Microsoft Excel

Put this experience to work on your problem.

Much of our work never reaches the website. Book a call, tell us your sector and we will walk you through the engagements that map to yours.