编程 Python

Python实现简单HTML表格解析的方法

Posted in Python onJune 15, 2015

本文实例讲述了Python实现简单HTML表格解析的方法。分享给大家供大家参考。具体分析如下：

这里依赖libxml2dom，确保首先安装！导入到你的脚步并调用parse_tables() 函数。

1. source = a string containing the source code you can pass in just the table or the entire page code

2. headers = a list of ints OR a list of strings
If the headers are ints this is for tables with no header, just list the 0 based index of the rows in which you want to extract data.
If the headers are strings this is for tables with header columns (with the tags) it will pull the information from the specified columns

3. The 0 based index of the table in the source code. If there are multiple tables and the table you want to parse is the third table in the code then pass in the number 2 here

It will return a list of lists. each inner list will contain the parsed information.

具体代码如下：

#The goal of table parser is to get specific information from specific
#columns in a table.
#Input: source code from a typical website
#Arguments: a list of headers the user wants to return
#Output: A list of lists of the data in each row
import libxml2dom
def parse_tables(source, headers, table_index):
  """parse_tables(string source, list headers, table_index)
    headers may be a list of strings if the table has headers defined or
    headers may be a list of ints if no headers defined this will get data
    from the rows index.
    This method returns a list of lists
    """
  #Determine if the headers list is strings or ints and make sure they
  #are all the same type
  j = 0
  print 'Printing headers: ',headers
  #route to the correct function
  #if the header type is int
  if type(headers[0]) == type(1):
    #run no_header function
    return no_header(source, headers, table_index)
  #if the header type is string
  elif type(headers[0]) == type('a'):
    #run the header_given function
    return header_given(source, headers, table_index)
  else:
    #return none if the headers aren't correct
    return None
#This function takes in the source code of the whole page a string list of
#headers and the index number of the table on the page. It returns a list of
#lists with the scraped information
def header_given(source, headers, table_index):
  #initiate a list to hole the return list
  return_list = []
  #initiate a list to hold the index numbers of the data in the rows
  header_index = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source,html=1)
  #get the tables from the document
  tables = doc.getElementsByTagName('table')
  try:
    #try to get focue on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get a list of headers in the table
  table_headers = main_table.getElementsByTagName('th')
  #need a sentry value for the header loop
  loop_sentry = 0
  #loop through each header looking for matches
  for header in table_headers:
    #if the header is in the desired headers list 
    if header.textContent in headers:
      #add it to the header_index
      header_index.append(loop_sentry)
    #add one to the loop_sentry
    loop_sentry+=1
  #get the rows from the table
  rows = main_table.getElementsByTagName('tr')
  #sentry value detecting if the first row is being viewed
  row_sentry = 0
  #loop through the rows in the table, skipping the first row
  for row in rows:
    #if row_sentry is 0 this is our first row
    if row_sentry == 0:
      #make the row_sentry not 0
      row_sentry = 1337
      continue
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #iterate through all of the header index's
    for i in header_index:
      #append the cells text content to the cell_list
      cell_list.append(cells[i].textContent)
    #append the cell_list to the return_list
    return_list.append(cell_list)
  #return the return_list
  return return_list
#This function takes in the source code of the whole page an int list of
#headers indicating the index number of the needed item and the index number
#of the table on the page. It returns a list of lists with the scraped info
def no_header(source, headers, table_index):
  #initiate a list to hold the return list
  return_list = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source, html=1)
  #get the tables from document
  tables = doc.getElementsByTagName('table')
  try:
    #Try to get focus on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get all of the rows out of the main_table
  rows = main_table.getElementsByTagName('tr')
  #loop through each row
  for row in rows:
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #loop through the list of desired headers
    for i in headers:
      try:
        #try to add text from the cell into the cell_list
        cell_list.append(cells[i].textContent)
      except:
        #if there is an error usually an index error just continue
        continue
    #append the data scraped into the return_list    
    return_list.append(cell_list)
  #return the return list
  return return_list

希望本文所述对大家的Python程序设计有所帮助。

Python实现简单HTML表格解析的方法

- Author -

小卒过河

声明：登载此文出于传递更多信息之目的，并不意味着赞同其观点或证实其描述。

Python 相关文章推荐

python实现自动更换ip的方法

May 05 Python

Python自动登录126邮箱的方法

Jul 10 Python

Python 专题一函数的基础知识

Mar 16 Python

Python实现删除文件中含“指定内容”的行示例

Jun 09 Python

Python 基础教程之str和repr的详解

Aug 20 Python

django ajax json的实例代码

May 29 Python

详解python while 函数及while和for的区别

Sep 07 Python

Pandas:Series和DataFrame删除指定轴上数据的方法

Nov 10 Python

Python属性和内建属性实例解析

Jan 14 Python

解决numpy矩阵相减出现的负值自动转正值的问题

Jun 03 Python

微软开源最强Python自动化神器Playwright(不用写一行代码)

Jan 05 Python

Python OpenCV之常用滤波器使用详解

Apr 07 Python

Python判断Abundant Number的方法

Jun 15 #Python

Python计算一个文件里字数的方法

Jun 15 #Python

Python素数检测实例分析

Jun 15 #Python

Python计算三维矢量幅度的方法

Jun 15 #Python

Python栈类实例分析

Jun 15 #Python

Python实现股市信息下载的方法

Jun 15 #Python

给Python入门者的一些编程建议

Jun 15 #Python

You might like

【星际争霸1】人族1v7家ZBath

2020/03/04 星际争霸

php你的验证码安全码？

2007/01/02 PHP

PHP处理excel cvs表格的方法实例介绍

2013/05/13 PHP

PHP扩展程序实现守护进程

2015/04/16 PHP

使用PHP生成二维码的方法汇总

2015/07/22 PHP

Thinkphp批量更新数据的方法汇总

2016/06/29 PHP

thinkphp5.1 文件引入路径问题及注意事项

2018/06/13 PHP

php微信公众号开发之二级菜单

2018/10/20 PHP

IE下js调试工具Companion.JS

2010/10/15 Javascript

基于jquery的loading效果实现代码

2010/11/05 Javascript

解决Extjs上传图片无法预览的解决方法

2012/03/22 Javascript

JavaScript中为什么null==0为false而null大于=0为true(个人研究)

2013/09/16 Javascript

jquery统计复选框选中示例

2013/11/05 Javascript

根据配置文件加载js依赖模块

2014/12/29 Javascript

JS+JSP通过img标签调用实现静态页面访问次数统计的方法

2015/12/14 Javascript

Web Uploader文件上传插件使用详解

2016/05/10 Javascript

Bootstrap框架动态生成Web页面文章内目录的方法

2016/05/12 Javascript

BootStrap table表格插件自适应固定表头(超好用)

2016/08/24 Javascript

vue实现ajax滚动下拉加载，同时具有loading效果(推荐)

2017/01/11 Javascript

干货!教大家如何选择Vue和React

2017/03/13 Javascript

JS开发中百度地图+城市联动实现实时触发查询地址功能

2017/04/13 Javascript

清空元素html("") innerHTML="" 与 empty()的区别和应用(推荐)

2017/08/14 Javascript

vue.js模仿京东省市区三级联动的选择组件实例代码

2017/11/22 Javascript

vue与bootstrap实现简单用户信息添加删除功能

2019/02/15 Javascript

JS实现的定时器展示简单秒表、页面弹框及跳转操作完整示例

2020/01/26 Javascript

使用Python获取Linux系统的各种信息

2014/07/10 Python

Python实现时钟显示效果思路详解

2018/04/11 Python

使用Python正则表达式操作文本数据的方法

2019/05/14 Python

美国珠宝店：Helzberg Diamonds

2018/10/24 全球购物

打架检讨书400字

2014/01/17 职场文书

小学班长竞选演讲稿

2014/04/24 职场文书

年终奖发放方案

2014/06/02 职场文书

读后感作文评语

2014/12/25 职场文书

升学宴来宾致辞

2015/07/27 职场文书

《给予树》教学反思

2016/03/03 职场文书

go设置多个GOPATH的方式

2021/05/05 Golang