Proceedings of the XMO Industrial Seminar 2026: Excellence in Manufacturing and Operations

Keywords

Large Language Model (LLM), LLM Reliability Assessment, Semantic Evaluation, Machine Monitoring

Tracks

MULTIFUNCTIONAL AND RESILIENT DESIGNS FOR MANUFACTURING

DOI

10.5703/1288284318673

Abstract

Large Language Models (LLMs) have recently shown strong performance across various domains, including manufacturing, where they are increasingly used with multi-agent systems for semi-autonomous tasks such as real-time machine monitoring and operator support. However, their reliability remains a concern due to hallucination, where outputs may appear plausible but contain incorrect or inconsistent information. To address this issue, this study proposes a semantic evaluation methodology for assessing the reliability of LLM-generated content in machining  environments. The approach is based on the embedding model BGE-M3 and focuses on capturing semantic consistency between machine observations and LLM-generated operational recommendations by comparing them against reference condition sentences and polarity anchors, classifying each monitoring sentence as Consistent, Contradiction, or Borderline. For evaluation, a set of realistic yet deliberately structured example sentences is constructed based on CNC machining scenarios. These examples are organized into standardized categories, allowing a more controlled analysis of semantic consistency and contextual coherence in LLM outputs. The results indicate that semantic analysis can reveal inconsistencies and subtle contextual deviations that may not be immediately apparent through manual inspection. This suggests that the proposed methodology can serve as a complementary tool for improving the reliability of LLM-based systems in manufacturing applications. Overall, this work highlights the promise of semantic approaches in supporting more trustworthy human–machine interaction, particularly in tasks such as anomaly interpretation and decision support during machining operations.

Share

COinS
 

A Semantic Evaluation Approach for Improving the Reliability of LLM Outputs in Manufacturing

Large Language Models (LLMs) have recently shown strong performance across various domains, including manufacturing, where they are increasingly used with multi-agent systems for semi-autonomous tasks such as real-time machine monitoring and operator support. However, their reliability remains a concern due to hallucination, where outputs may appear plausible but contain incorrect or inconsistent information. To address this issue, this study proposes a semantic evaluation methodology for assessing the reliability of LLM-generated content in machining  environments. The approach is based on the embedding model BGE-M3 and focuses on capturing semantic consistency between machine observations and LLM-generated operational recommendations by comparing them against reference condition sentences and polarity anchors, classifying each monitoring sentence as Consistent, Contradiction, or Borderline. For evaluation, a set of realistic yet deliberately structured example sentences is constructed based on CNC machining scenarios. These examples are organized into standardized categories, allowing a more controlled analysis of semantic consistency and contextual coherence in LLM outputs. The results indicate that semantic analysis can reveal inconsistencies and subtle contextual deviations that may not be immediately apparent through manual inspection. This suggests that the proposed methodology can serve as a complementary tool for improving the reliability of LLM-based systems in manufacturing applications. Overall, this work highlights the promise of semantic approaches in supporting more trustworthy human–machine interaction, particularly in tasks such as anomaly interpretation and decision support during machining operations.