250x250
Notice
Recent Posts
Recent Comments
Link
| 일 | 월 | 화 | 수 | 목 | 금 | 토 |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | 3 | 4 | 5 | 6 | 7 | 8 |
| 9 | 10 | 11 | 12 | 13 | 14 | 15 |
| 16 | 17 | 18 | 19 | 20 | 21 | 22 |
| 23 | 24 | 25 | 26 | 27 | 28 | 29 |
| 30 | 31 |
Tags
- Memory
- GPU
- 연결
- 교차검증
- oracle
- 머신러닝
- psql
- Jupyter
- sqldeveloper
- cpu
- Docker image
- Linux
- 도커
- Python
- 파이썬
- 쿼리
- postgre
- 오라클
- 복구
- docker
- GridSearchCV
- psycopg2
- 시계열
- pgadmin
- jupyternotebook
- 도커이미지
- TensorFlow
- LOG
- SQL
- 리눅스
Archives
- Today
- Total
areum
[NLP] 형태소 분석하기(feat. KoNLPy) 본문
728x90
Okt이용하여 형태소 분석해 보기 ( 형태소 빈도 측정)
1. 기본세팅
from collections import Counter
import numpy as np
import pandas as pd
from tqdm import tqdm
# konlpy
from konlpy.tag import Okt
okt = Okt()
# 데이터 불러오기
df = pd.read_csv('/content/ratings_train.txt',sep='\t')
# 데이터 양이 너무 많아 200행까지만 실행해보았습니다.
df=df.iloc[0:200,0:2]
2. Okt이용하여 형태소 나눠주기
sentences = []
for cp in tqdm(df.document.dropna()):
sentences.append(okt.pos(cp))
3. 단어 분류하기
words = []
for sentence in tqdm(sentences):
for word, tag in sentence:
if tag in ['Noun']:
words.append(word)
counts = Counter(words)
tags = counts.most_common(10)
tags

4. dataframe으로 만들어주기
df3 = pd.DataFrame(tags, columns=['단어', '빈도'])
df3

'Programming > NLP' 카테고리의 다른 글
| [NLP] 텍스트 카테고리 분류하는 방법 (0) | 2023.03.29 |
|---|---|
| [NLP] 텍스트 Emoji(이모티콘) 제거하는 방법 (1) | 2023.03.29 |