Iranian Journal of Information Processing and Management

Iranian Journal of Information Processing and Management

Automatic Intent Classification in a Task-Oriented Chatbot in Education

Document Type : Original Article

Authors
1 PhD Candidate; Faculty of Linguistics; Institute for Humanities and Cultural Studies
2 PhD in Computational Linguistics; Associate Professor; Faculty of Linguistics; Institute for Humanities and Cultural Studies
3 PhD in Linguistics; Associate Professor; Faculty of Linguistics; Institute for Humanities and Cultural Studies
10.22034/jipm.2026.2079807.2154
Abstract
Natural language constitutes a fundamental medium through which humans perform social actions, construct interpersonal relations, and negotiate cultural identities. Far beyond the mere transmission of propositional content, linguistic interaction embodies a structured system of practices—such as turn-taking, adjacency pairing, sequencing, and repair—which collectively enable coordinated social conduct. Since the emergence of Conversation Analysis in the 1960s, these interactional structures have been systematically examined, while Speech Act Theory, articulated by Austin and subsequently advanced by Searle, has underscored the intrinsically action-oriented nature of utterances. Within this theoretical landscape, the identification of a speaker’s communicative intent is a foundational requirement for meaningful interpretation. The advent of artificial intelligence and natural language processing has motivated efforts to computationally replicate this interpretive capacity within task-oriented dialogue systems, where intent detection serves as a core operational component.
Responding to the scarcity of annotated resources in Persian, the present study develops and evaluates a dedicated intent detection system for a Persian language-learning chatbot. To this end, a novel corpus comprising 3,607 sentences was constructed across four semantic relation types—synonymy, antonym, collocation, and idioms—derived from a systematic content analysis of authoritative Persian pedagogical materials and subsequently validated by domain experts. The dataset was thoroughly preprocessed, normalized, and annotated using the IOB scheme to capture the internal semantic structure of utterances. To model lexical meaning numerically, two static embedding methods grounded in distributional semantics, Word2Vec and FastText, were employed, generating 300-dimensional vector representations for all tokens.
Six machine learning algorithms—Support Vector Machine, K-Nearest Neighbors, Random Forest, Decision Tree, Single-Layer Perceptron, and Multi-Layer Perceptron—were trained and assessed using five-fold cross-validation. The experimental results demonstrate that model performance is substantially shaped by the choice of embedding method. FastText, owing to its subword-level encoding, facilitated markedly improved semantic discrimination and resulted in superior performance across most algorithms. The Multi-Layer Perceptron combined with FastText achieved the highest accuracy (99.45%), indicating its strong capacity to model non-linear semantic patterns in Persian. Conversely, Word2Vec exhibited weaker separability for closely related semantic categories, though it yielded strong outcomes in conjunction with distance-based models such as KNN.
Keywords
Subjects

  • Receive Date 04 December 2025
  • Revise Date 29 April 2026
  • Accept Date 06 June 2026