Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Kangping Yin

CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

Jul 06, 2021

Mosha Chen, Chuanqi Tan, Zhen Bi, Xiaozhuan Liang, Lei Li, Ningyu Zhang, Xin Shang, Kangping Yin, Jian Xu, Fei Huang(+13 more)

Figure 1 for CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

Figure 2 for CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

Figure 3 for CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

Figure 4 for CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

Abstract:Artificial Intelligence (AI), along with the recent progress in biomedical language understanding, is gradually changing medical practice. With the development of biomedical language understanding benchmarks, AI applications are widely used in the medical field. However, most benchmarks are limited to English, which makes it challenging to replicate many of the successes in English for other languages. To facilitate research in this direction, we collect real-world biomedical data and present the first Chinese Biomedical Language Understanding Evaluation (CBLUE) benchmark: a collection of natural language understanding tasks including named entity recognition, information extraction, clinical diagnosis normalization, single-sentence/sentence-pair classification, and an associated online platform for model evaluation, comparison, and analysis. To establish evaluation on these tasks, we report empirical results with the current 11 pre-trained Chinese models, and experimental results show that state-of-the-art neural models perform by far worse than the human ceiling. Our benchmark is released at \url{https://tianchi.aliyun.com/dataset/dataDetail?dataId=95414&lang=en-us}.

Via

Access Paper or Ask Questions

Conceptualized Representation Learning for Chinese Biomedical Text Mining

Aug 25, 2020

Ningyu Zhang, Qianghuai Jia, Kangping Yin, Liang Dong, Feng Gao, Nengwei Hua

Figure 1 for Conceptualized Representation Learning for Chinese Biomedical Text Mining

Figure 2 for Conceptualized Representation Learning for Chinese Biomedical Text Mining

Figure 3 for Conceptualized Representation Learning for Chinese Biomedical Text Mining

Figure 4 for Conceptualized Representation Learning for Chinese Biomedical Text Mining

Abstract:Biomedical text mining is becoming increasingly important as the number of biomedical documents and web data rapidly grows. Recently, word representation models such as BERT has gained popularity among researchers. However, it is difficult to estimate their performance on datasets containing biomedical texts as the word distributions of general and biomedical corpora are quite different. Moreover, the medical domain has long-tail concepts and terminologies that are difficult to be learned via language models. For the Chinese biomedical text, it is more difficult due to its complex structure and the variety of phrase combinations. In this paper, we investigate how the recently introduced pre-trained language model BERT can be adapted for Chinese biomedical corpora and propose a novel conceptualized representation learning approach. We also release a new Chinese Biomedical Language Understanding Evaluation benchmark (\textbf{ChineseBLUE}). We examine the effectiveness of Chinese pre-trained models: BERT, BERT-wwm, RoBERTa, and our approach. Experimental results on the benchmark show that our approach could bring significant gain. We release the pre-trained model on GitHub: https://github.com/alibaba-research/ChineseBLUE.

* WSDM2020 Health Day

Via

Access Paper or Ask Questions