html2article—基于文本密度的html2article实现[golag]Istallgo get -u -v github.com/sudy-li/html2articlePerformace
avg3.2msperarticle,accuracy>=98%(对比其他开源实现,可能是目前最快的html2article实现,我们测试的数据集约3kw来自于微信公众号,各大类中文科技媒体历史文章,目前能达到98%以上准确率)
Examples参考examples from_url.go
package maiimport ("github.com/sudy-li/html2article")fuc mai() {article, err := html2article.FromUrl("https://www.leiphoe.com/ews/201602/DsiQtR6c1jCu7iwA.html")if err != il {paic(err)}pritl("article title is =>", article.Title)pritl("article publishtime is =>", article.Publishtime)pritl("article cotet is =>", article.Cotet)}Algorithm参考论文
Java实现
评论