Skip to content

Add paste-into-document@pzim-devdata - #770

Open
pzim-devdata wants to merge 19 commits into
linuxmint:masterfrom
pzim-devdata:paste-into-document
Open

pzim-devdata wants to merge 19 commits into
linuxmint:masterfrom
pzim-devdata:paste-into-document

Conversation

@pzim-devdata

Copy link
Copy Markdown
Contributor

Paste clipboard content into new documents with format selection

  • Paste rich formatted text (HTML, bold, colors, tables preserved)
  • Paste images from clipboard or copied files
  • 11 output formats (ODT, PDF, TXT, MD, JSON, YAML, CSV, PY, SH, ODS, ODG)
  • Smart filename suggestion based on content
  • GTK dialog for filename and format selection
  • Document created without auto-opening
  • Includes 21 language translations
  • Distinct from existing "paste-into-file" (text only)

Paste clipboard content into new documents with format selection

- Paste rich formatted text (HTML, bold, colors, tables preserved)
- Paste images from clipboard or copied files
- 11 output formats (ODT, PDF, TXT, MD, JSON, YAML, CSV, PY, SH, ODS, ODG)
- Smart filename suggestion based on content
- GTK dialog for filename and format selection
- Document created without auto-opening
- Includes 21 language translations
- Distinct from existing "paste-into-file" (text only)
@pzim-devdata

pzim-devdata commented Jan 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Description

Paste clipboard content directly into a new document from Nemo's context menu with multiple format options. Unlike the existing "paste-into-file" action which only creates plain text files, this action creates formatted LibreOffice documents and offers 11 different output formats.

Features

  • Rich text support: Preserves HTML formatting (bold, italic, colors, tables, lists, etc.)
  • Image support: Paste images from clipboard or copied image files
  • Multiple formats: 11 output formats available
    • Text/Code: TXT, Markdown, JSON, YAML, CSV, Python, Shell script
    • Documents: LibreOffice Writer (ODT), PDF, Calc (ODS), Draw (ODG)
  • Smart filename: Auto-suggests filename based on clipboard content
    • Copied file → original filename
    • Text/HTML → first 3-5 words
    • Image → image_YYYYMMDD_HHMMSS
  • Format selector: GTK dialog for choosing output format
  • No auto-open: Creates document without opening it
  • Current location: Document created where user right-clicked

Why this action is useful

The existing "paste-into-file" action only creates plain text files without formatting. This action fills the gap by:

  • Preserving rich text formatting when pasting from web browsers, word processors, etc.
  • Supporting images (both clipboard and file URIs)
  • Offering multiple professional document formats
  • Providing LibreOffice document creation directly from Nemo

Distinction from existing action

  • Existing: "paste-into-file" → plain text only (.txt)
  • This action: "Paste into Document (.odt)" → rich formatting + 11 formats

The (.odt) in the name clearly distinguishes it from the text-only action.

Technical Details

  • Python 3 with GTK 3 for dialogs
  • LibreOffice headless mode for document conversion
  • HTML clipboard support for rich formatting
  • Base64 encoding for images
  • Automatic executable permission for .sh files
  • Two-step conversion for complex formats (HTML → ODT → PDF)

Dependencies

  • python3 (pre-installed on most systems)
  • libreoffice (for document conversion)

Testing

Tested on Debian Cinnamon with:

  • Rich formatted text from Firefox, Chrome, LibreOffice Writer
  • Images from GIMP, screenshot tool, copied files
  • All 11 output formats verified
  • Long filenames and special characters handled correctly

Screenshot

screenshot-auto_1769726190

Thank you to the Linux Mint team!

- Add HTML with formatting (.html) option to preserve web page structure
- Implement intelligent HTML text extraction for non-HTML formats
- Text formats (.txt, .md, .py, etc.) now automatically strip HTML tags
- Document formats (.odt, .pdf) extract clean text for readability
- Enhance filename suggestion to filter out HTML/CSS keywords
- Add strip_html_tags() function for clean text extraction
- Modify get_clipboard_content() to handle HTML based on selected format

This allows users to:
- Copy from web pages and get clean, readable text in most formats
- Preserve original HTML structure when specifically choosing .html format
- Get better default filenames when copying formatted web content
- Add HTML with formatting (.html) option to preserve web page structure
- Implement intelligent HTML text extraction for non-HTML formats
- Text formats (.txt, .md, .py, etc.) now automatically strip HTML tags
- Document formats (.odt, .pdf) extract clean text for readability
- Enhance filename suggestion to filter out HTML/CSS keywords
- Add strip_html_tags() function for clean text extraction
- Modify get_clipboard_content() to handle HTML based on selected format

This allows users to:
- Copy from web pages and get clean, readable text in most formats
- Preserve original HTML structure when specifically choosing .html format
- Get better default filenames when copying formatted web content
@pzim-devdata

Copy link
Copy Markdown
Contributor Author

Add HTML format option and improve web content handling

  • Add HTML with formatting (.html) option to preserve web page structure
  • Implement intelligent HTML text extraction for non-HTML formats
  • Text formats (.txt, .md, .py, etc.) now automatically strip HTML tags
  • Document formats (.odt, .pdf) extract clean text for readability
  • Enhance filename suggestion to filter out HTML/CSS keywords
  • Add strip_html_tags() function for clean text extraction
  • Modify get_clipboard_content() to handle HTML based on selected format

This allows users to:

  • Copy from web pages and get clean, readable text in most formats
  • Preserve original HTML structure when specifically choosing .html format
  • Get better default filenames when copying formatted web content

@pzim-devdata

pzim-devdata commented Feb 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Improve default filename generation with "most repeated word" strategy

Enhanced the slugify() function to generate smarter default filenames by finding the most important word and extracting its natural context.

The script now uses a "most repeated word" approach:

  1. Most repeated word: Counts word occurrences (excluding common words), finds the most frequent word, and extracts the phrase leading up to its first occurrence (max 4-5 words before it)
  2. HTML <title> tag: Detects title from web pages
  3. HTML headings: Extracts from <h1>, <h2>, <h3> tags
  4. Link text: Uses first meaningful link as fallback

Examples:

  • "je suis heureux" → je_suis_heureux
  • "...elle me rend heureux... je suis heureux..." → elle_me_rend_heureux
  • "...corrigé la version. Le problème était... Le problème venait..." → le_problème
  • <h1>Guide d'installation Python</h1> → guide_installation_python

Also adds full Unicode support for accented characters (é, è, à, ç) and keeps ALL words in the final filename (including articles) for natural, readable phrases.

This approach automatically identifies the document's key topic rather than taking arbitrary first words.

Comment thread paste-into-document@pzim-devdata/files/paste-into-document@pzim-devdata/check.py Outdated
Title changed from 'Paste into Document (.odt)' to 'Paste into Document':
the action no longer targets a single format. All po files updated
accordingly (msgid + msgstr).

New features:
- 'No extension added' as the default dropdown entry: the filename
  gets no extension unless the user selects a format or types one
- 14 additional extensions: js, ts, css, php, c, cpp, java, rb, xml,
  svg, ini, toml, sql, tsv, plus docx and rtf for documents
- Two distinct HTML modes: 'HTML file from code' saves copied source
  code exactly as-is (plain text target), 'Raw HTML' keeps the
  clipboard text/html target of a copied web page

Bug fix:
- Clipboard content is now snapshotted once, before any dialog opens.
  Previously, copying text while naming the file (e.g. pasting the
  name itself) replaced the original content in the final document.
@pzim-devdata

Copy link
Copy Markdown
Contributor Author

Rename action and improve format handling

Title changed from 'Paste into Document (.odt)' to 'Paste into Document':
the action no longer targets a single format. All po files updated
accordingly (msgid + msgstr).

New features:

  • 'No extension added' as the default dropdown entry: the filename
    gets no extension unless the user selects a format or types one
  • 14 additional extensions: js, ts, css, php, c, cpp, java, rb, xml,
    svg, ini, toml, sql, tsv, plus docx and rtf for documents
  • Two distinct HTML modes: 'HTML file from code' saves copied source
    code exactly as-is (plain text target), 'Raw HTML' keeps the
    clipboard text/html target of a copied web page

Bug fix:

  • Clipboard content is now snapshotted once, before any dialog opens.
    Previously, copying text while naming the file (e.g. pasting the
    name itself) replaced the original content in the final document.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants