all options
buster  ] [  bullseye  ] [  bookworm  ] [  trixie  ] [  sid  ]
[ Source: htmlcxx  ]

Package: libhtmlcxx3v5 (0.87-4)

Links for libhtmlcxx3v5

Screenshot

Debian Resources:

Download Source Package htmlcxx:

Maintainers:

External Resources:

Similar packages:

simple HTML parser library for C++

htmlcxx is a simple non-validating CSS1 and HTML parser for C++. Although there are several other html parsers available, htmlcxx has some characteristics that make it unique:

 * STL like navigation of DOM tree, using excellent tree.hh library from
   Kasper Peeters
 * It is possible to reproduce exactly, character by character, the original
   document from the parse tree
 * Bundled CSS parser
 * Optional parsing of attributes
 * C++ code that looks like C++ (not so true anymore)
 * Offsets of tags/elements in the original document are stored in the nodes
   of the DOM tree

The parsing politics of htmlcxx were created trying to mimic Mozilla Firefox (http://www.mozilla.org) behavior. So you should expect parse trees similar to those create by Firefox. However, differently from Firefox, htmlcxx does not insert non-existent stuff in your html. Therefore, serializing the DOM tree gives exactly the same bytes contained in the original HTML document.

Other Packages Related to libhtmlcxx3v5

  • depends
  • recommends
  • suggests
  • enhances

Download libhtmlcxx3v5

Download for all available architectures
Architecture Package Size Installed Size Files
arm64 28.4 kB150.0 kB [list of files]