MARATTO

article · Statistical Journal of the IAOS

Automating enterprise industry classifications for official statistics: Leveraging text-based similarity measures

Abstract

Accurate industrial classification of firms forms the backbone of business surveys, economic policymaking, and international trade analysis. However, national statistics institutes (NSIs) worldwide grapple with the labor intensive manual assignment of International Standard Industrial Classification (ISIC) codes: a process prone to human error, inconsistent across regions, and particularly burdensome for developing economies. This study confronts these challenges by assessing performance of token-overlap (Jaccard), TF-IDF cosine similarity, edit-distance (fuzzy) and SBERT embeddings against human-coded ground truth in classifying firms. Using a dataset of 6588 firms, performance diverges sharply: SBERT attains Accuracy <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mo>=</mml:mo> </mml:math> 0.78 and Weighted <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mrow> <mml:mi mathvariant="normal">F</mml:mi> <mml:mn>1</mml:mn> </mml:mrow> <mml:mo>=</mml:mo> <mml:mn>0.78</mml:mn> </mml:math> (Cohen’s <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mi>κ</mml:mi> <mml:mo>≈</mml:mo> <mml:mn>0.75</mml:mn> </mml:math> ), while surface methods lag (Fuzzy: Accuracy 0.43; Cosine: 0.31; Jaccard: 0.26). Statistical tests confirms these differences (Cochran’s ( <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mi>Q</mml:mi> <mml:mo>=</mml:mo> <mml:mn>8320.81</mml:mn> </mml:math> ) with <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mi>p</mml:mi> <mml:mo>&lt;</mml:mo> <mml:mn>0.001</mml:mn> </mml:math> ) and inter-method agreement is only fair ( <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:msub> <mml:mi>κ</mml:mi> <mml:mtext>Fleiss</mml:mtext> </mml:msub> <mml:mo>≈</mml:mo> <mml:mn>0.270</mml:mn> </mml:math> ), motivating a class-level diagnostic approach. Using confusion matrices and Haberman adjusted residuals we expose systematic off-diagonal confusions (notably between manufacturing, professional/service and certain retail/wholesale categories) and identify classes with strong, automatable diagonals versus sparse or ambiguous tails that require human coding.

Research topics

  • Economic and Technological Innovation
  • Data Quality and Management
  • Time Series Analysis and Forecasting

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1177/18747655261433544

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.