kamike.collect 网络爬虫_开源项目-程序员客栈

开源地址
https://github.com/hubinix/kamike.collect授权协议
未知

AotherSimpleCrawler又一个网络爬虫，可以支持代理服务器的翻墙爬取。

1.数据存在mysql当中。

2.使用时，先修改web-if/cofig.ii的数据链接相关信息，主要是数据库名和用户名和密码

3.然后访问https://127.0.0.1/fetch/istall链接，自动创建数据库表

4.修改src\java\c\exihua\fetch中的RestServlet.java文件：

FetchIst.getIstace().ruig=true; Fetch fetch = ew Fetch(); fetch.setUrl("https://www.washigtopost.com/"); fetch.setDepth(3); RegexRule regexRule = ew RegexRule(); regexRule.addNegative(".*#.*"); regexRule.addNegative(".*pg.*"); regexRule.addNegative(".*jpg.*"); regexRule.addNegative(".*gif.*"); regexRule.addNegative(".*js.*"); regexRule.addNegative(".*css.*"); regexRule.addPositive(".*php.*"); regexRule.addPositive(".*html.*"); regexRule.addPositive(".*htm.*"); Fetcher fetcher = ew Fetcher(fetch); fetcher.setProxyAuth(true); fetcher.setRegexRule(regexRule); List<Fetcher> fetchers = ew ArrayList<>(); fetchers.add(fetcher); FetchUtils.start(fetchers); 将其配置为需要的参数，然后访问https://127.0.0.1/fetch/fetch启动爬取代理的配置在Fetch.java文件中： protected it status;protected boolea resumable = false;protected RegexRule regexRule = ew RegexRule();protected ArrayList<Strig> seeds = ew ArrayList<Strig>();protected Fetch fetch;protected Strig proxyUrl="127.0.0.1";protected it proxyPort=4444;protected Strig proxyUserame="hkg";protected Strig proxyPassword="deis";protected boolea proxyAuth=false;

5.访问https://127.0.0.1/fetch/susped可以停止爬取

Another Simple Crawler 又一个网络爬虫，可以支持代理服务器的翻墙爬取。 1.数据存在mysql当中。 2.使用时，先修改web-inf/config.ini的数据链接相关信...

声明：本文仅代表作者观点，不代表本站立场。如果侵犯到您的合法权益，请联系我们删除侵权资源！如果遇到资源链接失效，请您通过评论或工单的方式通知管理员。未经允许，不得转载，本站所有资源文章禁止商业使用运营!

下载安装【程序员客栈】APP

实时对接需求、及时收发消息、丰富的开放项目需求、随时随地查看项目状态

前往安装

kamike.collect 网络爬虫开源项目

技术信息

作品详情

功能介绍

重点城市程序员兼职推荐

重点岗位程序员兼职推荐