Home Table of Contents

25 August 2026, Volume 10 Issue 7-8
    

  • Select all
    |
  • Wu Xuan, Li Guangjian, Pan Jiali, Wang Chuhan
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 1-14. https://doi.org/10.11925/infotech.2096-3467.2025.0777
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the limitations of existing research on U.S. export controls—namely insufficient targeting, limited correlation accuracy, and narrow analytical dimensions—this paper explores a method for assessing the Sino-U.S. competitive landscape in key core technologies within the export control context. [Methods] By integrating multi-source data including strategic technology lists, export control lists, and patent data, this study employs large language models (LLMs) to automate the association among strategic technologies, controlled items, and patents. An indicator system based on the T-GCP (Technology Gap-Controllability-Potential) three-dimensional assessment model is designed to evaluate the Sino-U.S. competitive landscape in key core technologies. [Results] An empirical study in the quantum technology field establishes precise linkages between 13 distinct Export Control Classification Number (ECCN) items across quantum computing, communication, and sensing domains and their corresponding patent data. Using a technology competition risk matrix, the study carries out comprehensive competitive situation assessment and risk zoning. [Limitations] The current assessment indicator system requires further expansion and enrichment, and its capacity for predictive analysis of potential future control domains remains insufficient. [Conclusions] The proposed methodology enables systematic assessment of the Sino-U.S. competitive landscape and technology risk levels in key core technology fields, providing a decision support tool for responding to technology controls and formulating breakthrough strategies.

  • Cai Yiran, Hu Zhengyin, Chen Wenjie, Xu Haiyun, Han Tao
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 15-28. https://doi.org/10.11925/infotech.2096-3467.2025.0718
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This paper proposes a scientific literature dataset construction framework for AI4S to improve the AI readiness of scientific literature data. [Methods] We first analyze the core data requirements of AI4S and develop a four-level data architecture: L1 raw literature data, L2 multimodal decomposition data, L3 synthetic reasoning data, and L4 task application data. A complete dataset construction workflow is then designed, covering compliant data acquisition, intelligent annotation, synthetic reasoning, and application and evaluation. Large language models (LLMs) are integrated with a human-in-the-loop strategy to improve both dataset construction efficiency and data quality. An organic solar cell (OSC) domain dataset is constructed to validate the proposed framework. [Results] The resulting AI4S-oriented OSC scientific literature dataset contains 2,711 raw publications, 31,428 multimodal decomposed records, and 5,343 synthetic reasoning records, for a total of 39,482 data objects. Compared with conventional scientific literature datasets, the proposed dataset demonstrates superior structural completeness, semantic richness, and better applicability to diverse AI4S applications. It can effectively support downstream tasks such as organic solar cell experiment recommendation and can be extended to other AI4S application domains. [Limitations] The proposed framework still has limitations in cross-modal semantic association of complex formulas and multidimensional charts. In addition, the quality of synthetic data, and dataset construction efficiency can be further enhanced, while the feasibility and effectiveness of fine-grained data provenance require further validation. [Conclusions] This study presents a reusable framework for constructing AI-ready scientific literature datasets. The proposed methodology enhances the structured and intelligent governance of such data, thereby strengthening the capability of scientific literature data to support AI4S applications.

  • Tian Xuecan, Li Changwang, Liu Chen, Deng Zeyu, Mao Jin
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 29-41. https://doi.org/10.11925/infotech.2096-3467.2025.0863
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To characterize the anomalous fluctuations of innovation activities within frontier technology fields amidst global competition, providing a reference for identifying vulnerable points within the technological system. [Methods] Leveraging longitudinal patent application data, this study employed multiple time-series forecasting and anomaly detection models to quantify anomalous performance across technological fields. The EScore was utilized to evaluate their frontier status. Through cross-dimensional analysis, anomaly profiles were constructed for these frontier technology fields. [Results] The optimal performance of different models varies across technology fields. Overall, China’s technological innovation system exhibits “macro-level stability punctuated by localized drastic fluctuations”. Most detected anomalies are negative, with abrupt shifts typically demonstrating a 2-to-3-year lag. A general negative correlation exists between a field’s frontier level and its anomaly severity; highly frontier technology fields are more susceptible to negative fluctuations, although technology fields such as the Internet of Things (IoT) and computational chemistry demonstrate significant resilience. [Limitations] Innovation activities were solely measured by patent application volume, and the frontier assessment relied on a single metric. Future research should incorporate more comprehensive indicators. [Conclusions] Constructing anomaly profiles for frontier technologies based on the dual dimensions of frontier level and anomaly severity effectively reveals systemic vulnerabilities and potential breakthrough trajectories for technological innovation in a highly competitive international landscape.

  • Ding Shengchun, Gong Jingze, Qin Tianyun
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 42-55. https://doi.org/10.11925/infotech.2096-3467.2025.0848
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] Aiming at the challenges of diverse types, complex interactive behaviors, and poor reproducibility of cognitive subjects in the context of cognitive warfare, this study introduces large language model (LLM) technology to construct a multi-agent interactive simulation model based on real-world data. [Methods] An interactive simulation model was constructed based on data from the Al-Ahli Hospital explosion incident. First, deep learning methods were employed to extract the attribute distributions and behavioral probability characteristics of nine typical categories of cognitive subjects. Second, LLMs were utilized to generate 10,000 agent instances with individual heterogeneity. Finally, these agents were integrated into the NetLogo platform to conduct interactive behavior simulations. The validity of the model was verified from two dimensions: attribute distribution consistency and behavioral pattern differentiation. [Results] The interactive simulation model accurately characterizes the influence differences among cognitive subjects across different social strata. Differentiated interactions conforming to a normal distribution emerged during the simulations, effectively overcoming the limitations of rigid rule-setting and insufficient sample representativeness inherent in traditional simulation approaches. [Limitations] The current model primarily focuses on fitting and reproducing behavioral probabilities, and does not yet support dynamic cognitive evolution based on real-time semantic interactions during simulation operation, resulting in insufficient capacity for evaluating the effectiveness of in-depth semantic confrontation. [Conclusions] The interactive simulation model constructed in this study can effectively reproduce behavioral responses following information reception, thereby validating the feasibility of the technical approach that combines LLM-generated agents with NetLogo-based simulation. It provides quantifiable and reproducible simulation-based deduction support for cognitive warfare operations.

  • Liu Xiumin, Hu Maodi, Song Donghuan, Sun Xi, Zou Dong, Yuan Zhixiang
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 56-67. https://doi.org/10.11925/infotech.2096-3467.2025.0849
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To improve the performance of large language models in extracting knowledge objects from technological patents, this study addresses the limitations of few-shot prompt method and the insufficient alignment between demonstrations and target tasks. We propose a method for knowledge object extraction from technology patent via dynamic demonstration prompting. [Methods] Demonstration selection is formulated as a demonstration-guided gain (DGG) prediction problem. A task-aware cross-encoder ranking model is constructed to capture the deep semantic interaction between target patent sentences and candidate demonstrations. A two-stage retrieval-reranking framework is then employed to retrieve candidates and dynamically select demonstrations with higher demonstration-guided gains. [Results] Experiments on a Chinese genomics patent dataset show that the dynamic demonstration selection method achieves an F1 score of 64.60% in knowledge object extraction, outperforming the baseline model. These results verify the effectiveness of the DGG model in improving the quality of dynamically selected demonstrations. [Limitations] The experimental validation of this study is mainly based on Chinese patent texts in the genomics domain, with other technology domains and text types not yet covered; the applicability of the proposed method to other domains and text types warrants further investigation. [Conclusions] Validated through experiments on genomics patents, the proposed method enhances demonstration-task alignment and improves the performance of LLM-based knowledge object extraction from technology patents.

  • Li Jinhao, Zhao Yuxiang, Zhao Yanke, Zhu Qinghua
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 68-79. https://doi.org/10.11925/infotech.2096-3467.2025.0827
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This study aims to explore the characteristics and triggers of users’ nostalgic emotions in nostalgic videos based on user comments. [Methods] Over 20,000 user comments from 40 nostalgic-themed videos on Bilibili were subjected to computational grounded analysis. Pattern detection, refinement, and confirmation were conducted using the Qwen large language model and prompt engineering. [Results] Five primary nostalgic elements—characters, events, time, place, and objects—were extracted from video comments. Three influencing factors triggering nostalgia were identified: sensory processing, recommendation mechanisms, and social interaction, alongside two sociocultural attributes. [Limitations] Findings derived from computational grounding may overlook insights in user comments due to biases inherent in unsupervised classification. [Conclusions] This study deepens the understanding of core dimensions in nostalgic content creation on social media and provides reference points for designing user experiences in nostalgic videos.

  • Zhou Wenhao, Lin Hongxi, Zhang Zhiwei, Li Yawen
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 80-92. https://doi.org/10.11925/infotech.2096-3467.2025.0782
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This study aims to overcome the constraints of the conventional single-layer network perspective, which falls short of revealing the complex mechanisms underlying technological innovation. A multi-layer network frame is developed, encompassing corporate collaboration networks, knowledge combination networks, and R&D team networks. Using the biopharmaceutical manufacturing sector as an empirical case, the study systematically identifies the key driving factors and cross-layer interaction mechanisms that influence firms' technological innovation performance. [Methods] The study integrates multi-source databases, including China Stock Market & Accounting Research Database, Wind and Innojoy patent to construct a three-layer network embedding framework comprising inter-firm collaboration networks, knowledge combination networks, and R&D team networks. Relevant network characteristics and firm-level baseline variables are extracted accordingly. A random forest regression model is then applied to systematically examine the determinants and their relative importance for technological innovation performance. In addition, the SHAP method is employed to quantify the marginal contribution and nonlinear interactive effects of individual features, thereby establishing an interpretable analytical framework for the multi-layer network influence mechanisms. [Results] The full-variable model incorporating multi-layer network embeddedness features outperforms the benchmark model that includes only baseline firm characteristics in explaining technological innovation performance, thereby validating the effectiveness of a multi-source relational data fusion strategy in predictive modeling of complex social systems. Feature importance analysis based on SHAP values reveals that, among the baseline features, R&D investment exhibits the highest marginal contribution. Within the multi-layer network features, relationship depth in the corporate cooperation network, closeness centrality in the knowledge combination network, and cohesion in the R&D team network are respectively identified as the most explanatory predictors at each network layer. [Limitations] The generalizability of the findings is constrained by the industry-specific focus. In addition, the dynamic evolution of network characteristics has not been sufficiently accounted for, and there remains considerable scope for further exploration and in-depth investigation into the interaction mechanisms among multiple network layers. [Conclusions] This study not only theoretically advances beyond the confines of traditional innovation studies and expands the application scope of multi-layer network theory but also offers an operable analytical framework and methodological support for building data-driven innovation management systems in complex network environments.

  • Yan Qiang, Leng Jidong, Yi Lanli, Jiang Lidan
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 93-104. https://doi.org/10.11925/infotech.2096-3467.2025.0775
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To reveal the mechanisms underlying public cognitive and emotional responses during the diffusion of disruptive technologies and to explain the underlying logic of their social acceptance. [Methods] A three-stage response model of “technology perception-emotional expression-attitudinal stance” was constructed. Taking Apollo Go as the research case, more than 60,000 user comments from Bilibili, Douyin, and Xiaohongshu were analyzed using semantic analysis, emotion recognition, and attitudinal stance classification to compare cross-platform differences in public technology perception, emotional expression, and attitudinal stance. [Results] Public technology perception exhibited a multidimensional structure, emotional expression showed significant platform heterogeneity, and attitudinal stances were clearly divided between support for the technology and institutional concerns. By shaping the generation and expression of emotion, platforms played a key moderating role in the formation of public attitudes. [Limitations] The case analysis was mainly based on the Chinese-language social media context, and the applicability of the findings to other linguistic and cultural contexts remains to be verified. [Conclusions] The study demonstrates the interconnected mechanism among perception, emotion, and attitude in the diffusion of disruptive technologies, offers a new perspective for constructing public technology acceptance models, and provides methodological references for AI-driven public opinion analysis and scenario-based decision-making.

  • Wang Haoyu, Zhou Yulin, Huang Ruizhang, Qin Yongbin
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 105-114. https://doi.org/10.11925/infotech.2096-3467.2025.0743
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the issues of insufficient legal knowledge integration and poor judgment compliance in existing sentencing prediction models for multi-defendant cases, this paper proposes a Knowledge-Aided Sentencing Prediction (KASP) method that integrates legal-doctrinal constraints with knowledge-driven strategies. [Methods] This study employs large language models (LLMs) to decompose case facts, analyze corresponding charges and legal provisions, and extract basic sentencing intervals as structured legal priors. A consistency fusion mechanism is adopted to integrate the legal prior knowledge into the training process of a lightweight prediction model, thereby achieving knowledge-driven collaborative optimization. [Results] Experimental results on the CMDL-small dataset demonstrate that, compared with the optimal baseline DeepSeek-R1-14B, KASP achieves improvements of 5.44 and 4.18 percentage points in accuracy and F1-score, respectively, for the sentencing prediction task, exhibiting more stable performance in complex multi-defendant scenarios. [Limitations] This study primarily focuses on knowledge modeling for extracting basic sentencing intervals from legal-doctrinal constraints and discretionary sentencing factors. It does not address more complex sentencing rules, such as combined punishment for multiple crimes or concurrence of legal provisions. [Conclusions] Incorporating structured legal prior knowledge can effectively enhance both the predictive performance and legal compliance of sentencing prediction models in complex cases.

  • Zou Limin, Xu Xiaoyue, Pan Weipeng, Li Wan, Li Xiuting
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 115-128. https://doi.org/10.11925/infotech.2096-3467.2025.0654
    Abstract ( ) Download PDF ( ) HTML ( )   Knowledge map   Save

    [Objective] This study investigates how firms’ structural positions in multiplex relational networks influence their decisions regarding digital transformation, thereby providing a micro-level explanation of corporate transformation behavior from a network-relations perspective. [Methods] Using Chinese A-share-listed property service firms from 2011 to 2022 as the research sample, this study applies complex network theory to construct shareholding, industry, and interpersonal networks. It quantifies firms’ structural positions and influence within these networks and examines the conditions under which interfirm connections affect digital transformation performance. [Results] The multiplex relational characteristics of firms affect their digital transformation processes to varying degrees. Customer, shareholding, industry, and interpersonal connections all significantly influence firms’ digital transformation decisions. Transaction-connection and interpersonal-connection strategies strengthen the positive effects of digital transformation; however, excessive emphasis on these strategies may inhibit the effectiveness of digital transformation. [Limitations] The sample is primarily concentrated in the property service industry, and the relational networks are constructed mainly from data covering specific dimensions. [Conclusions] From a complex network perspective, this study enriches the theoretical understanding of the drivers of digital transformation, confirms the importance of firms’ network positions, and provides decision-making guidance for firms seeking to leverage network resources to accelerate their transformation.

  • Ao Yuxuan, Wang Hao, Zhou Shu, Bu Wenru
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 129-140. https://doi.org/10.11925/infotech.2096-3467.2025.0755
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the limitations of existing poetry-to-image generation techniques in extracting poetic semantics and enhancing the artistic expressiveness of generated images, this study proposes a poetry-to-image generation method that balances semantic alignment with aesthetic quality. [Methods] The method comprises three modules: deep semantic understanding, multimodal imagery space construction, and multi-stage collaborative generation. A pre-trained language model and multi-task learning are used to extract emotion, poetic imagery, and rhetorical features and align them with visual representations, followed by image generation in three stages. [Results] On the self-constructed Poetic Visions dataset, compared with the best-performing baseline (GPT-4o+DALL·E 2), the proposed method improves the Inception Score (IS) from 25.32 to 26.87, reduces the Fréchet Inception Distance (FID) from 15.75 to 14.98, raises the CLIP Score from 0.68 to 0.72, and increases the human evaluation score from 3.3 to 3.7. [Limitations] The multi-stage process depends on early-stage composition outputs, and deviations may propagate to subsequent stages. Fine-grained control and generation stability remain limited for long or highly abstract poems. [Conclusions] Semantic guidance and collaborative generation effectively improve poem-image semantic alignment and artistic expression, demonstrating application potential in digital humanities and creative scenarios.

  • Hao Xiping, Yu Chuanming, Zhang Dianyuan, Fu Xueqing
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 141-143. https://doi.org/10.11925/infotech.2096-3467.2025.0774
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the high computational cost and limited accuracy improvements associated with prompt-enhanced methods for multiple-choice machine reading comprehension (MCRC) based on large language models (LLMs), this study proposes Prompt-MCRC, an prompt-enhanced multiple-choice machine reading comprehension with large language models. [Methods] The proposed model employs a large language model as the teacher model and uses chain-of-thought (CoT) prompting to guide a local student model in extracting and integrating the teacher model’s reasoning patterns and auxiliary signals. A bidirectional matching mechanism is further introduced to strengthen fine-grained semantic associations. In addition, a Japanese multiple-choice machine reading comprehension dataset, JLPT-MC, is independently constructed based on the Japanese-Language Proficiency Test (JLPT). [Results] The proposed model consistently outperformed the standalone local-model baselines on the C3-M, RACE-M, JGLUE, JaQuAD, and JLPT-MC datasets. Compared with the best-performing large language model baselines, it achieved improvements of 0.83 percentage points in F1 score on JaQuAD and 2.89 percentage points in accuracy on JLPT-MC, respectively. [Limitations] Owing to the high cost of manual annotation, the self-constructed dataset is relatively small in scale, which limits the model’s cross-lingual generalization capability to some extent. [Conclusions] By integrating chain-of-thought prompting and a bidirectional matching mechanism, the proposed model improves the reasoning performance of local models on multiple-choice machine reading comprehension tasks and provides a new perspective for multilingual research in this field.

  • Zhang Jiacheng, Liu Zheli, Xiao Guangwen, Nie Lihai, Wang Yongchang, Shi Liang, Jin Meihong
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 154-167. https://doi.org/10.11925/infotech.2096-3467.2025.0850
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the fragmentation of existing value-alignment evaluation frameworks for large language models (LLMs), their insufficient coverage of values with Chinese characteristics, the scarcity of in-depth evaluation data, and the limitations of current evaluation methods, this study develops methods and techniques for evaluating value alignment in LLMs. [Methods] This study proposes an integrated methodological framework combining value rules, evaluation data, and intelligent technologies. Within this framework, a three-dimensional evaluation system encompassing capabilities, tasks, and metrics is designed. Data are collected, augmented, and annotated to construct an in-depth evaluation dataset. Finally, a value-alignment evaluation model is trained by integrating pretrained models, instruction fine-tuning, and expert feedback. [Results] The resulting evaluation model achieves an accuracy of 98.57%, enabling automated assessment of value alignment in LLMs. The empirical results indicate that Chinese-developed models generally exhibit higher levels of alignment than overseas models. Nevertheless, they still commonly suffer from insufficient training data related to Chinese revolutionary culture, factual inaccuracies and hallucinatory misinformation, attenuation of ideological content, over-censorship, and limited capacity for dynamic adaptation. [Limitations] This study focuses primarily on text-based LLMs and therefore has limited applicability to multimodal models. Moreover, the evaluation results are presented using only three levels—high, medium, and low—leaving room for improvement in interpretability. [Conclusions] This study contributes to the development of a value-alignment governance framework with Chinese characteristics, supports the sound development of LLMs within a safe, trustworthy, and controllable framework, and provides technical support for effectively incorporating China’s mainstream values into economic development and social governance.

  • Han Mingxing, Lin Litao, Ou Shiyan, Xu Liwei
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 168-178. https://doi.org/10.11925/infotech.2096-3467.2025.0834
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This study proposes a social bot detection model that integrates large language model representations with a mixture-of-experts model to address the limited ability to distinguish fine-grained semantic differences in textual content and fuse multi-source heterogeneous features. [Methods] The proposed model first uses LLM2Vec to embed user posts and capture fine-grained semantic differences. It then integrates six categories of features—user, social-network, temporal, content, sentiment, and AI-related features—and employs a routing network to dynamically assign them to the corresponding expert networks. [Results] Experimental results on two public datasets (TwiBot-20 and TwiBot-22) show that the proposed model achieves precisions of 0.8210 and 0.7756, respectively. Its overall performance surpasses that of the comparison models, thus validating the effectiveness of the proposed model. [Limitations] This study has not yet examined the impact of adversarial attacks on model robustness or thoroughly discussed the ethical risks that may arise from the misclassification of social bots. [Conclusions] LLM2Vec facilitates the extraction of semantic information from user posts, while the mixture-of-experts model enables differentiated modeling and the effective fusion of multi-source heterogeneous features, thereby improving the performance of social bot detection.

  • Li Zhiwen, Zhang Le, Zhao Tianming, Che Chao
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 179-191. https://doi.org/10.11925/infotech.2096-3467.2025.0840
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the frequent hallucination and limited long-context understanding exhibited by existing approaches to comprehensive health examination report generation, this study proposes a Gated Large Language Model (LLM) and Adaptive Retrieval-Augmented Generation (RAG) approach. [Methods] The proposed approach comprises three components: a report-generation LLM, a gating module, and an adaptive RAG module. The report-generation LLM is fine-tuned on a dataset preprocessed through BERT-based entity extraction and manual semantic chunking. The gating module consists of a Forget Gate and a Memory Gate. At the input stage, the Forget Gate uses the LLM’s prior knowledge to identify and remove subqueries that have no impact on the global query. At the output stage, the Memory Gate integrates the question-answer pairs produced during report generation and the outputs of automatic retrieval augmentation, using Chain-of-Thought (CoT) reasoning to validate and consolidate diagnostic conclusions. Furthermore, the adaptive RAG module selects external knowledge bases with heterogeneous structures according to different examination results, thereby enhancing retrieval capability and inference efficiency. [Results] On a self-built dataset of personal comprehensive health examination reports, the proposed approach achieved BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L scores of 0.47, 0.65, 0.51, and 0.67, respectively, outperforming the comparison models on most metrics. [Limitations] The proposed approach still exhibits limitations in handling unstructured medical text, in preventing potential irreversible information loss caused by hard pruning, and in updating knowledge when the knowledge base is static. [Conclusions] Experimental results show that the proposed approach significantly improves the long-context understanding and medical proficiency of billion-parameter-scale LLMs in comprehensive health examination report generation, resulting in higher-quality reports.

  • Yang Renbiao, Cao Gaohui
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 192-204. https://doi.org/10.11925/infotech.2096-3467.2025.0769
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This study focuses on the realization pathways of misinformation herd immunity and aims to explore effective approaches for misinformation governance. [Methods] A Realization Pathways of Herd Immunity to Misinformation Model (RP-MHIM) incorporating two evolutionary mechanisms—inoculation-based intervention and natural infection—was developed. By introducing multiple intervention variables, including inoculation frequency, timing, and intensity, the model systematically simulates and analyzes immunity outcomes under different pathways. [Results] In terms of immunity speed, the inoculation pathway achieves herd immunity at approximately t=150 iterations, whereas the natural infection pathway requires around 200 iterations. From the perspective of pathway robustness, the inoculation strategy exhibits significantly stronger resistance to disturbances than the natural infection strategy. [Limitations] The model still relies on simplified assumptions concerning individual behavior, platform response mechanisms, and network structures, and its strategy effects have not been empirically validated using real-world data. [Conclusions] The vaccination-based pathway consistently outperforms the natural infection pathway across multiple dimensions, demonstrating higher intervention efficiency and a faster immunity formation process.

  • Peng Mingyang, Gao Yan, Lai Yuqiao
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 205-216. https://doi.org/10.11925/infotech.2096-3467.2025.0780
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This study proposes an ordinal-aware hierarchical fusion network (OAFHN) to address two limitations of existing harmful meme detection methods: insufficient modeling of the ordinal progression of harmfulness levels, and the application of symmetric penalties to different types of misclassification. [Methods] This study first designs an ordinal-aware and false-positive penalty loss (OPP-Loss), which reformulates classification as ordinal regression and applies an asymmetric penalty to false positives. It then constructs a hierarchical multipath fusion network that takes semantic paraphrases generated by a vision-language model as knowledge input, and combines coarse-grained fusion, semantically modulated attention, and low-rank bilinear pooling to model features at multiple granularities. [Results] OAFHN achieves F1 scores of 83.46% and 88.39% on the Harm-C and Harm-P datasets, exceeding the strongest baseline MOMENTA by 0.66 and 0.13 percentage points, respectively. In the ablation study, replacing OPP-Loss with cross-entropy loss reduces the F1 score on Harm-C by 8.27 percentage points, and removing or replacing the other components also causes performance degradation to varying degrees. [Limitations] The false-positive penalty factor requires manual tuning, and the fixed ordinal mapping does not fully capture the heterogeneity within the “partially harmful” class. [Conclusions] Under the experimental settings used in this study, ordinal-aware optimization and hierarchical multipath fusion can improve the performance of harmful meme detection, but the false-positive penalty factor and the ordinal mapping scheme still require further optimization.

  • Li Xue, Sun Lijuan, Gao Yutong, Nan Guoshun, Li Gaohu, Wu Xu
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 217-226. https://doi.org/10.11925/infotech.2096-3467.2025.0752
    Abstract ( ) Download PDF ( ) HTML ( )   Knowledge map   Save

    [Objective] This paper proposes an event evolutionary graph-based retrieval-augmented network negative event detection model (EEG-RAG) to address the issues of inaccurate semantic understanding and deficient perception of event evolutionary relationships in real-time online negative event detection by large language models. [Methods] The model is built upon a self-constructed event evolutionary graph of online negative events, which serves as the knowledge backbone. It mitigates semantic bias via an event classification learning mechanism and incorporates event evolutionary logic to implement retrieval augmentation, endowing the underlying large language model with precise online negative event detection and interpretable decision-making capabilities. [Results] Compared with baseline models and models fine-tuned on domain-specific data, the proposed model achieves a significant improvement in semantic understanding accuracy. The retriever achieves a precision of over 91% and a 36.8 percentage point improvement in recall. For online negative event detection, the model achieves an accuracy gain of 6.8 percentage points and a recall gain of 8.5 percentage points. [Limitations] This study is confined to the text and does not encompass the multimodal information of online negative events. [Conclusions] The proposed model effectively enhances the semantic understanding of similar events, improves the recall rate of associated events, and reduces misclassification rates. This provides reliable technical support for the real-time monitoring of online negative events.

  • Ma Yanzhou, Luo Yun, Wu Shengyi, Zhu Qi
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 227-237. https://doi.org/10.11925/infotech.2096-3467.2025.0823
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This study focuses on topic-related sarcasm detection on Chinese social media. Current large language models (LLMs) are often overly sensitive, misclassifying strong opinions or purely negative emotions as sarcasm, which undermines their effectiveness and accuracy in this area. [Methods] We design a dual-path reasoning framework. On the inductive path, the model retrieves analogous cases with theoretical explanations to facilitate analogy-based reasoning. The deductive path involves constructing a two-stage judgment framework grounded in Impoliteness Theory and designing two-stage prompts that guide the model from sentiment filtering to intention analysis. Finally, a dual-path evidence fusion module integrates evidence from both paths to generate the final judgment. [Results] Comparative experiments on the ToSarcasm benchmark demonstrate that our approach outperforms established baselines. Using DeepSeek-V3.1 as the base model, our framework achieves an F1 score of 83.25% and a macro F1 score of 76.45%, surpassing the best traditional baseline by 10.44 and 14.63 percentage points, respectively. [Limitations] This research was primarily conducted within a Chinese context, so its cross-lingual generalizability requires further validation. While the deductive framework enhances interpretability, its decision-making process fundamentally relies on the internal mechanisms of the LLM, which are not fully transparent. [Conclusions] The proposed theory-guided, dual-path reasoning framework effectively enhances the robustness and balance of LLMs in identifying topic-related sarcasm, providing a new approach to tackling complex pragmatic reasoning tasks.

  • Zhou Jian, Lyu Lucheng, Xu Jiayuan, Zheng Lili, Zhao Yajuan
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 238-250. https://doi.org/10.11925/infotech.2096-3467.2025.0839
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] This paper proposes a large language model (LLM)-based framework for automatic patent technology monitoring to improve the efficiency of technical monitoring in intelligence services. This framework enables periodic and automated monitoring of field-related patents. [Methods] The proposed LLM-based framework consists of the following steps: periodic patent data collection, semantic vector-based coarse screening, LLM classification-based fine screening, multi-dimensional indicator ranking, and automatic bulletin generation. Experiments on patent semantic vector similarity, patent classification and bulletin generation were conducted using the Qwen3-Embedding-8B model and the GPT-OSS-120B model, with the AI for Science (AI4S) field as a case study. Performance was evaluated in terms of classification precision, classification stability, and bulletin hallucination, through manual annotation and multi-model cross-validation. [Results] In the AI4S patent classification task, the framework achieved a macro-average classification precision of 97.81%, with a precision of 96% for the top 200 high-ranked patents. The LLM-based fine screening demonstrated strong stability. The generated bulletin exhibited a low level of hallucination in both multi-LLM evaluations and expert evaluations. [Limitations] The empirical study focuses solely on the AI4S field, and further evaluations across additional frontier domains are required. Meanwhile, hallucination in LLMs may persist so that future intelligence work should continue to rely on human-AI collaboration to ensure reliability. [Conclusions] The proposed LLM-based framework for automatic patent technology monitoring significantly enhances the efficiency and intelligence of identifying technologies and generating bulletins. It provides valuable support for automated scientific and technological intelligence, research trend analysis, and strategic decision-making.

  • Qiu Jingwen, Wang Hao, Yang Simin, Yao Tianchen, Tan Yuyao
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 251-262. https://doi.org/10.11925/infotech.2096-3467.2025.0858
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the challenge of achieving large-scale and objective profiling of historical figures, this paper proposes a historical figure profiling framework integrating event-level sentiment analysis, aiming to comprehensively and objectively analyze the traits of historical figures. [Methods] Using the Shiji (《史记》, Records of the Grand Historian) as the experimental corpus, this study first develops a Integrating Semantic and Structural Information Recognition Model (ISSI-RM) to extract figure-behavior units and construct an event dataset. An Attention-Based Generative Multi-Label Classification Model (ABG-MLC) model is then developed to automatically assign labels to historical events. Prompt learning is subsequently employed to recognize the sentiment of historical events. Finally, figure profiling is performed across three dimensions: event classification, the alignment between event-level sentiment and evaluative sentiment, and the temporal evolution of event-level sentiment. [Results] The ISSI-RM identifies approximately three groups of figure-behavior units per sentence on average and effectively reduces the ambiguity and errors in entity coreference resolution. The ABG-MLC model achieves an F1 score of 0.80 in multi-label classification, an absolute improvement of 0.14 over the baseline models. The complete pipeline, which first performs event extraction and classification and then applies prompt learning for sentiment recognition, achieves the best recognition performance with an F1 score of 0.86. Empirical results show that the label combination “political competence + military competence” accounts for the largest proportion among multi-label events and appears predominantly among the founding emperors in the Annals and the strategists in the Biographies. The consistency between event-level sentiment and evaluative sentiment differs markedly between figures in the Annals and those in the Biographies; in the samples where the two sentiments are highly consistent, positive sentiment predominates. The temporal event-sentiment trajectories of the historical figures can be clustered into four typical evolutionary patterns. [Limitations] This study draws exclusively on the Shiji as its experimental corpus, which limits the diversity of textual sources. [Conclusions] By leveraging digital technologies, this study integrates event-level sentiment analysis into historical figure profiling, addressing the limitations of traditional historical figure research in scalability and objectivity. The framework offers a fresh perspective for the study of historical figures, as well as practical guidance for the deep-driven exploration of historical texts and the conversion of historical materials into cultural resources.

  • Wang Nan, Wang Juan, Liu Yaowen, Pan Jie, Xia Yixue, Xian Tingyu
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 263-278. https://doi.org/10.11925/infotech.2096-3467.2025.0742
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the reliance on single visual artifacts and the limited robustness to interference in existing AIGC image attribution methods, a semantic-guided multimodal attribution model, named MIRAGE, is proposed. [Methods] To bridge the semantic gap between multimodal features, the model first introduces a semantic mapping module that translates quantitative fingerprints into natural language. Building upon this, MIRAGE adjusts the interaction logic of the cross-attention layer by utilizing the generated semantic text as an active query, guiding the model to dynamically focus on critical artifact evidence within the deep feature space. [Results] MIRAGE achieves F1 scores of 98.4% and 69.6% on the WILD and DRAGON datasets, respectively. Quantitative analysis demonstrates that compared to the unimodal visual baseline, MIRAGE improves the F1 score by 5.9 and 11.3 percentage points, respectively. Furthermore, compared to the simple feature concatenation baseline, it yields improvements of 3.1 and 6.4 percentage points in complex scenarios, respectively. [Limitations] The rule-based semantic generation constrains adaptive reasoning capabilities; additionally, the discriminability among homologous models with highly similar architectures requires further enhancement. [Conclusions] The semantic-guided active multimodal fusion model effectively integrates orthogonal evidence, offering an effective pathway to enhance the robustness of AIGC image attribution in complex scenarios.

  • Zhang Jingyuan, Yang Lei, Liu Zhaiyi
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 279-292. https://doi.org/10.11925/infotech.2096-3467.2025.0773
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the issues of high computational complexity and limited multi-scale feature extraction capability in Transformer-based models, which make it difficult to simultaneously capture local details and global contextual information, this study proposes a Multiscale Lightweight Attention Network (MLA-Net) for depression recognition. [Methods] MLA-Net adopts a lightweight Transformer architecture and integrates a global dual-pooling attention mechanism to extract video features while preserving global information. Subsequently, an attention mechanism is employed to model long-range spatiotemporal dependencies. Multi-scale feature extraction is further incorporated to capture information at different scales, and a cross-fusion strategy is utilized to enhance feature representation. [Results] Experimental results on a real-world depression dataset demonstrate that MLA-Net achieves a mean absolute error (MAE) of 4.90 and a root mean square error (RMSE) of 6.88, outperforming the comparison methods and validating its effectiveness and rationality. [Limitations] The proposed method only analyzes a single modality, namely facial expressions, without integrating other modalities such as speech, text, or physiological signals. [Conclusions] Through the collaborative effects of global dual-pooling attention, multi-scale feature extraction, and cross-feature fusion, MLA-Net significantly improves the performance of depression recognition.

  • Liu Tiantian, Peng Fang, Zhu Tianyou, Yang Chao
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 293-301. https://doi.org/10.11925/infotech.2096-3467.2025.0679
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address the low logical-form and execution accuracy of natural-language-to-SQL models caused by insufficient access to target-database schema information in real-world deployments, this study investigates the integration of retrieval-augmented generation into the SQL generation process. [Methods] A large language model named SQLGPT is proposed to translate users’ natural-language queries into SQL statements. SQLGPT first performs semantic similarity search to retrieve database schema information relevant to a given query. It then combines the retrieved schema information with in-context learning examples to dynamically construct prompts, which guide the large language model in generating the corresponding SQL statements. [Results] Experiments on the WikiSQL show that SQLGPT achieves a logical-form accuracy of 86.1% and an execution accuracy of 91.9%, outperforming BRIDGE by 0.4 and 0.8 percentage points, respectively. The model also supports multi-turn interaction, demonstrating its effectiveness in handling successive database queries. [Limitations] Although SQLGPT performs well on WikiSQL, its evaluation is limited to a single general-purpose dataset. Its robustness and generalizability have yet to be fully validated in scenarios involving irregular schema naming, complex multi-turn interactions, and domain-specific queries. [Conclusions] By integrating semantic similarity-based schema retrieval with dynamic prompt construction, SQLGPT provides an effective and accurate approach to natural-language-to-SQL generation and help reduce the barriers to database access for users without specialized technical expertise.

  • Wei Wei, Yang Long, Ding Shuangying
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 302-318. https://doi.org/10.11925/infotech.2096-3467.2025.0847
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] Focusing on knowledge extraction of policy synergy in the cultural, tourism, and sports (CTS) industries, this study designs an intelligent knowledge extraction pipeline based on large language models (LLMs) for extracting policy synergy features, evaluating synergy levels, and identifying synergy evolution mechanisms. [Methods] By employing social network analysis, an intelligent policy instrument recognition algorithm, and an LLM-optimized BERTopic model, this study systematically extracts the evolutionary features of central policy texts and proposes an LLM-based intelligent policy synergy evaluation framework (PMC_LLMs) to achieve automated assessment of policy synergy at the local level. [Results] From 418 central policy documents, 15 categories of policy instruments and 6 core themes are identified, revealing the internal synergistic evolution pathways of the policy system. Furthermore, based on PMC_LLMs, the synergy levels of 119 local specialized policies are evaluated, and the distribution patterns and evolutionary dynamics of 46 sub-indicators are identified. [Limitations] This study focuses on the internal synergy mechanisms within the policy system and has not yet extended to examining the actual impact of external pathways on industry evolution. [Conclusions] This study validates the effectiveness and applicability of LLMs in intelligent policy synergy evaluation, providing theoretical support and practical reference for policy intelligence research and the optimization of CTS industry policies.

  • Jiang Zhan, Yu Chuanming, Lyu Guangliu
    Data Analysis and Knowledge Discovery. 2026, 10(7-8): 319-331. https://doi.org/10.11925/infotech.2096-3467.2025.0762
    Abstract ( ) Download PDF ( ) HTML   Knowledge map   Save

    [Objective] To address data sparsity and suboptimal recognition performance in patent named entity recognition (NER), this study explores the application of large language models (LLMs) to this task. [Methods] We propose a ChatGLM4.5-enhanced data augmentation and multi-task joint optimization method. The method uses prompt templates to guide ChatGLM4.5 in generating augmented patent samples, encodes both original and augmented samples with BERT, constructs positive and negative entity pairs for contrastive learning, and integrates CRF sequence labeling with cross-entropy and CRF losses for joint optimization. [Results] On the TFH-2020 and CPIE datasets, the proposed method outperforms the overall best-performing baseline by 2.81 and 1.74 percentage points in F1 score, respectively. [Limitations] The experiments are limited to a small number of publicly available patent datasets and have not yet covered additional technical domains or multilingual patent text scenarios. [Conclusions] The LLM-based augmentation and joint optimization strategy improves patent entity recognition performance and provides a feasible approach for applying LLMs to patent information extraction.