================================================================================= KCC150, KCCq28, KCC940 -- Korean Contemporary Corpus of Written Sentences Total 732 million words (48,878,948 sentences) ================================================================================= KCC(Korean Contemporary Corpus) -- raw sentences of the Korean langugae 1) KCC150 --150,705,457 words (11,961,347 sentences) Sentences with quotes are not included. 2) KCCq28 -- 28,782,776 words (1,337,721 sentences with double qoutes) All the sentences include a quote. 3) KCC940 -- 93,210,332 words (6,263,454 sentences) All the sentences are no more than 30 words. 4) KCC460 -- about 460 million words (29,316,426 sentences) All the sentences are no more than 30 words. 5) KCC text corpus for Word2Vec word embedding of the above KCC corpus. [Download] Korean word embedding model Seung-Shik Kang, Ph.D Professor at Kookmin University Email: nlpkang AT g m a i l . c o m =================================================================================
Download one of the following(UTF8 encoding).
- KCC150_Korean_sentences_UTF8.txt.gz -- "UTF8" encoded file
- KCCq28_Korean_sentences_UTF8_v2.zip -- "UTF8" encoded file (0xA1A1 -> 0x20)
- KCC940_Korean_sentences_UTF8_V2.txt.gz -- "UTF8" encoded file (merge Grapheme to word)
- KCC460 -- no "UTF8" encoded file. Download KCC460_EUCKR.txt.gz and convert to UTF8 by iconv library!
- http://203.246.112.71/sskang/kcc/KCC460.txt.gz -- EUCKR encoding
Download iconv --> $ iconv -c -f utf-8 -t -cp949 KCC460_EUCKR.txt > KCC460_utf8.txt
Download one of the following text files for word embedding(Word2Vec) --> EUCKR encoded files
These files are automatically created by KLT2000 Korean morphological analyzer. See below for the details.
C> index2018.exe -c test.txt output.txt https://cafe.naver.com/nlpkang/3 https://cafe.naver.com/nlpk/278- KCC150_Korean_sentences for Word2Vec embedding ("EUCKR" encoded file)
================================================================================ 대용량 파일 다운로드 문제로 인하여 다운받기 어려운 경우들이 발생하는 경우에... ================================================================================ 1) KCC150 -- 1억5천만 어절을 분할하여 다운로드 KCC150 11,961,347문장(1억5천만 어절)은 100만 문장(라인)씩 12개 파일로 분할. 아래 12개 파일을 순서대로 결합하면 KCC150_Korean_sentences_UTF8.txt와 동일함(단, encoding은 EUCKR) - KCC150_K01.txt.gz - KCC150_K02.txt.gz - KCC150_K03.txt.gz - KCC150_K04.txt.gz - KCC150_K05.txt.gz - KCC150_K06.txt.gz - KCC150_K07.txt.gz - KCC150_K08.txt.gz - KCC150_K09.txt.gz - KCC150_K10.txt.gz - KCC150_K11.txt.gz - KCC150_K12.txt.gz 2) KCCq28 -- 2,878만 어절을 분할하여 다운로드 큰따옴표 포함된 1,331,721문장(2,878만 어절)은 20만 문장(라인)씩 7개 파일로 분할. 아래 7개 파일을 순서대로 결합하면 KCCq28_Korean_sentences_UTF8.txt와 동일함(단, encoding은 EUCKR) - KCCq28_Q01.txt.gz - KCCq28_Q02.txt.gz - KCCq28_Q03.txt.gz - KCCq28_Q04.txt.gz - KCCq28_Q05.txt.gz - KCCq28_Q06.txt.gz - KCCq28_Q07.txt.gz