Python实现简单HTML表格解析的方法


Posted in Python onJune 15, 2015

本文实例讲述了Python实现简单HTML表格解析的方法。分享给大家供大家参考。具体分析如下:

这里依赖libxml2dom,确保首先安装!导入到你的脚步并调用parse_tables() 函数。

1. source = a string containing the source code you can pass in just the table or the entire page code

2. headers = a list of ints OR a list of strings
If the headers are ints this is for tables with no header, just list the 0 based index of the rows in which you want to extract data.
If the headers are strings this is for tables with header columns (with the tags) it will pull the information from the specified columns

3. The 0 based index of the table in the source code. If there are multiple tables and the table you want to parse is the third table in the code then pass in the number 2 here

It will return a list of lists. each inner list will contain the parsed information.

具体代码如下:

#The goal of table parser is to get specific information from specific
#columns in a table.
#Input: source code from a typical website
#Arguments: a list of headers the user wants to return
#Output: A list of lists of the data in each row
import libxml2dom
def parse_tables(source, headers, table_index):
  """parse_tables(string source, list headers, table_index)
    headers may be a list of strings if the table has headers defined or
    headers may be a list of ints if no headers defined this will get data
    from the rows index.
    This method returns a list of lists
    """
  #Determine if the headers list is strings or ints and make sure they
  #are all the same type
  j = 0
  print 'Printing headers: ',headers
  #route to the correct function
  #if the header type is int
  if type(headers[0]) == type(1):
    #run no_header function
    return no_header(source, headers, table_index)
  #if the header type is string
  elif type(headers[0]) == type('a'):
    #run the header_given function
    return header_given(source, headers, table_index)
  else:
    #return none if the headers aren't correct
    return None
#This function takes in the source code of the whole page a string list of
#headers and the index number of the table on the page. It returns a list of
#lists with the scraped information
def header_given(source, headers, table_index):
  #initiate a list to hole the return list
  return_list = []
  #initiate a list to hold the index numbers of the data in the rows
  header_index = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source,html=1)
  #get the tables from the document
  tables = doc.getElementsByTagName('table')
  try:
    #try to get focue on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get a list of headers in the table
  table_headers = main_table.getElementsByTagName('th')
  #need a sentry value for the header loop
  loop_sentry = 0
  #loop through each header looking for matches
  for header in table_headers:
    #if the header is in the desired headers list 
    if header.textContent in headers:
      #add it to the header_index
      header_index.append(loop_sentry)
    #add one to the loop_sentry
    loop_sentry+=1
  #get the rows from the table
  rows = main_table.getElementsByTagName('tr')
  #sentry value detecting if the first row is being viewed
  row_sentry = 0
  #loop through the rows in the table, skipping the first row
  for row in rows:
    #if row_sentry is 0 this is our first row
    if row_sentry == 0:
      #make the row_sentry not 0
      row_sentry = 1337
      continue
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #iterate through all of the header index's
    for i in header_index:
      #append the cells text content to the cell_list
      cell_list.append(cells[i].textContent)
    #append the cell_list to the return_list
    return_list.append(cell_list)
  #return the return_list
  return return_list
#This function takes in the source code of the whole page an int list of
#headers indicating the index number of the needed item and the index number
#of the table on the page. It returns a list of lists with the scraped info
def no_header(source, headers, table_index):
  #initiate a list to hold the return list
  return_list = []
  #get a document object out of the source code
  doc = libxml2dom.parseString(source, html=1)
  #get the tables from document
  tables = doc.getElementsByTagName('table')
  try:
    #Try to get focus on the desired table
    main_table = tables[table_index]
  except:
    #if the table doesn't exits then return an error
    return ['The table index was not found']
  #get all of the rows out of the main_table
  rows = main_table.getElementsByTagName('tr')
  #loop through each row
  for row in rows:
    #get all cells from the current row
    cells = row.getElementsByTagName('td')
    #initiate a list to append into the return_list
    cell_list = []
    #loop through the list of desired headers
    for i in headers:
      try:
        #try to add text from the cell into the cell_list
        cell_list.append(cells[i].textContent)
      except:
        #if there is an error usually an index error just continue
        continue
    #append the data scraped into the return_list    
    return_list.append(cell_list)
  #return the return list
  return return_list

希望本文所述对大家的Python程序设计有所帮助。

Python 相关文章推荐
python中单下划线_的常见用法总结
Jul 10 Python
Django实现支付宝付款和微信支付的示例代码
Jul 25 Python
python框架中flask知识点总结
Aug 17 Python
python 使用re.search()筛选后 选取部分结果的方法
Nov 28 Python
使用Python实现微信提醒备忘录功能
Dec 04 Python
Python Flask框架模板操作实例分析
May 03 Python
python3实现斐波那契数列(4种方法)
Jul 15 Python
Python交互式图形编程的实现
Jul 25 Python
Python操作列表常用方法实例小结【创建、遍历、统计、切片等】
Oct 25 Python
Python如何实现强制数据类型转换
Nov 22 Python
Python变量、数据类型、数据类型转换相关函数用法实例详解
Jan 09 Python
完美解决keras保存好的model不能成功加载问题
Jun 11 Python
Python判断Abundant Number的方法
Jun 15 #Python
Python计算一个文件里字数的方法
Jun 15 #Python
Python素数检测实例分析
Jun 15 #Python
Python计算三维矢量幅度的方法
Jun 15 #Python
Python栈类实例分析
Jun 15 #Python
Python实现股市信息下载的方法
Jun 15 #Python
给Python入门者的一些编程建议
Jun 15 #Python
You might like
php数组函数序列之end() - 移动数组内部指针到最后一个元素,并返回该元素的值
2011/10/31 PHP
完美解决PHP中的Cannot modify header information 问题
2013/08/12 PHP
php打造智能化的柱状图程序,用于报表等
2015/06/19 PHP
php使用crypt()函数进行加密
2017/06/08 PHP
IE网页js语法错误2行字符1、FF中正常的解决方法
2013/09/09 Javascript
js实现瀑布流的一种简单方法实例分享
2013/11/04 Javascript
js简单实现删除记录时的提示效果
2013/12/05 Javascript
基于jQuery全屏焦点图左右切换插件responsiveslides
2015/09/07 Javascript
jQuery实现选项卡功能(两种方法)
2017/03/08 Javascript
bootstrap suggest下拉框使用详解
2017/04/10 Javascript
基于jstree使用AJAX请求获取数据形成树
2017/08/29 Javascript
移动web开发之touch事件实例详解
2018/01/17 Javascript
Angular4 组件通讯方法大全(推荐)
2018/07/12 Javascript
简单的React SSR服务器渲染实现
2018/12/11 Javascript
小程序云开发之用户注册登录
2019/05/18 Javascript
jQuery 动画与停止动画效果实例详解
2020/05/19 jQuery
解决idea开发遇到javascript动态添加html元素时中文乱码的问题
2020/09/29 Javascript
[02:20]DOTA2英雄基础教程 黑暗贤者
2013/12/19 DOTA
Python实现字符串的逆序 C++字符串逆序算法
2020/05/28 Python
Python实现删除时保留特定文件夹和文件的示例
2018/04/27 Python
重写django的model下的objects模型管理器方式
2020/05/15 Python
理肤泉加拿大官网:La Roche-Posay加拿大
2018/07/06 全球购物
Coltorti Boutique官网:来自意大利的设计师品牌买手店
2018/11/09 全球购物
Prototype如何更新局部页面
2013/03/03 面试题
信息专业毕业生五年职业规划参考
2014/02/06 职场文书
团代会主持词
2014/04/02 职场文书
书香校园建设方案
2014/05/02 职场文书
职务任命书范本
2014/06/05 职场文书
2014新生大学四年计划书
2014/09/21 职场文书
第二批党的群众路线教育实践活动个人对照检查材料
2014/09/23 职场文书
领导干部“四风”问题批评与自我批评材料
2014/09/24 职场文书
2014年实习期工作总结
2014/11/27 职场文书
2015廉洁自律个人总结
2015/02/14 职场文书
试用期旷工辞退通知书
2015/04/17 职场文书
pytorch 两个GPU同时训练的解决方案
2021/06/01 Python
详解在OpenCV中如何使用图像像素
2022/03/03 Python