1 / 7

SVM based Spam Filtering in SEWM2007

SVM based Spam Filtering in SEWM2007. Pan Weike, Lu Guanzhong, Xu Congfu panweike@zju.edu.cn , oillgz@gmail.com , xucongfu@zju.edu.cn College of Computer Science, Zhejiang University March 11, 2007. Chinese Anti-spam Framework. Outline. Email Pre-processing Feature Extraction

hanna-riggs
Download Presentation

SVM based Spam Filtering in SEWM2007

An Image/Link below is provided (as is) to download presentation Download Policy: Content on the Website is provided to you AS IS for your information and personal use and may not be sold / licensed / shared on other websites without getting consent from its author. Content is provided to you AS IS for your information and personal use only. Download presentation by click this link. While downloading, if for some reason you are not able to download a presentation, the publisher may have deleted the file from their server. During download, if you can't get a presentation, the file might be deleted by the publisher.

E N D

Presentation Transcript


  1. SVM based Spam Filtering in SEWM2007 Pan Weike, Lu Guanzhong, Xu Congfu panweike@zju.edu.cn, oillgz@gmail.com, xucongfu@zju.edu.cn College of Computer Science, Zhejiang University March 11, 2007

  2. Chinese Anti-spam Framework

  3. Outline • Email Pre-processing • Feature Extraction • Support Vector Regression

  4. Email Pre-processing • Some problems: • An email may contain more than 2 charset types. • The charset information of some emails are missing. • An efficient approach to obtain the accurate charset information of each email is needed.

  5. Feature Extraction • Tokenization: Tianwang Chinese algorithm • http://net.pku.edu.cn/~webg/src/ChSeg/ • Without Feature Selection: TF,CHI,IG, etc. • VSM: TF*IDF, subject:body=3:1

  6. Support Vector Regression • SVR toolbox: libSVM http://www.csie.ntu.edu.tw/~cjlin/libsvm/

  7. Thanks for your attention!

More Related