PDFContentExtractor Class
Definition
Namespace: O2S.Components.PDF4NET.Content
Defines an extractor for various type of content from PDF pages.
Inheritance: Object → PDFContentExtractor
Constructor
| Name | Description |
|---|---|
| PDFContentExtractor(PDFPage) | Initializes a new PDFContentExtractor object. |
Properties
| Name | Description |
|---|---|
| Encodings | Gets the list of additional encodings used for extracting text from PDF files. |
Methods
| Name | Description |
|---|---|
| ExtractColorSpaces | Extracts the colorspaces that exist in the page resources. |
| ExtractContentStreamOperators | Extracts the page content stream as a list of graphic operators with their operands. |
| ExtractImages | Extracts the information related to the images displayed on the page. |
| ExtractOptionalContentGroup | Extracts the content of an optional content group. |
| ExtractText | Extracts the text from the PDF page. |
| ExtractTextLines | Extracts the text from the PDF page as a collection of PDFTextLine objects. |
| ExtractTextRuns | Extracts the text fragments from the PDF page. |
| ExtractVisualObjects | Extracts the page content as a list of visual objects. |
| ExtractWords | Extracts the text from the PDF page as a collection of words. |
| SearchText | Searches the page content for the specified text. |