Saturday, February 19, 2022

Limitations of Calibre and Calibre-Web as an epub book viewer

When viewing an epub book using Calibre, one's position in the book is indicated as percent, in the bottom right corner of the reader window. This is OK for a small book but it does not show fractions of a percent and therefore is the same for many pages of a longer book. It does have the virtue that it is independent of the size of the reader window. Calibre maintains a constant position when the window is resized: at least some of the text is common between the larger and smaller window.

The Calibre-Web reading window is worse: it doesn't indicate position in the book at all. This is made worse by the fact that if one resizes the reader window the reader jumps to a different position in the book. If one alternates between a small and large window, just flipping back and forth, position gradually moves to the beginning of the current chapter. There is no practical way to get back to the original position except to go back to the start of the chapter then scan forward to the desired location. This is very poor UX.

In the case of Calibre-Web, the behaviour is a result of using epubjs to display the book. I have made a small test app that uses default configuration and it exhibits the same faults: no indication of position and position jumps when window is resized.

The epubjs package has many options. There may be options to display position and to maintain position when window is resized, but I haven't found them yet. I tried capturing and repositioning after resize, but it is a bit of a nightmare of rapid-fire events and deciding when to save and when to restore position. There does not appear to be any option to disable the automatic redisplay on window size change and there is essentially no documentation of the implementation.

As with so much software these days, the documentation of epubjs provides the syntax of the API but very little to nothing regarding the semantics.

There are some clues about getting page information in issue 744. None of this defines what is meant by "page" but most of the comments imply that what is meant is: what is presented in the reader window.

Calibre-Web has an inbuilt table-of contents for epub books but if it is opened the right side of the content goes off-screen with no scrollbar. It is impossible to read the text with the table of contents open.

Calibre-Web allows bookmarks to be saved but the presentation of them is as a cryptic string like "epubcfi(/6/8[id_4]!/4/2/1:0)". This is a standard epub cfi but it is not human friendly. Given a list of these, how would a normal person know which refers to what? There is no way to annotate them. In contrast, a bookmark in Firefox can be edited to change the title, icon, etc.

The bookmarks persist between browser sessions.  And they persist through clearing cookies and site data. This suggests that they are stored on the server. But presumably not in the book itself as each user could have different bookmarks.

It appears that Calibre-Web has its own database, in addition to the Calibre database it accesses. Default on Linux is ~/.calibre-web/app.db. This includes table 'bookmark' with fields user_id, book_id, format and bookmark_key, the latter two beeing epub and an epubcfi string. So, bookmarks are per user, per book and persistent across devices and sessions.

Both epub2 and epub3 have support for page lists but various posts suggest that almost no epub2 and few epub3 readers actually use them.

The navigation file provides some guidance on navigation for epub3.

A book might be published in hard copy, with a particular layout of pages. The same book in epub format will break the text into different sections, depending on font, window size, etc. One might be interested in what page of the original hard copy publication is being viewed, regardless of how many reader windows it takes to view the complete page, or one might be interested in pages as determined by the reader window: one reader window full is one page.

Is it possible to position the reader window at an arbitrary position in the text? If so, then what 'page' is the reader at when the displayed text starts at the second character of the book? Or the third? etc.

If the definition of "page" depends on the size of the reader window, font, screen resolution for displaying images, etc. then page number is only relevant in the current reader window context. If one reads the same book on a different system, with a different reader size, with a different font size, etc. then the page numbers will all be different. "Page 237" will contain different text depending on all these (and probably other) factors.

Counting characters might be more consistent: current display starts at character "1073648 of 27634287", for example. But this isn't very human friendly.

Pages are a familiar concept in paper books, but what do they mean in an e-book?

How does one return to the same position in the book, when one re-opens it?

How does one refer to a part of the book in computer friendly terms (where character offset might be fine) and human friendly terms?

Is it possible to make a fixed page list that is independent of the reader window size, screen resolution, font size, etc.? 

There is a good discussion of epub3 page lists at epubsecrets. This includes use cases where consistency across different media formats, independent of the individual reader details, is useful.

The epub3 spec includes page-list nav element.

epub3 has support for fixed layout documents, with the introduction pointing out that by default epub3 documents are intended to adapt to the reader with reflow, etc.

Firefox / Librewolf unsigned add-ons

In short:

Enable unsigned add-ons by setting xpinstall.signatures.required to false in about:config.

Make sure the add-on manifest.json has an id, as in the following:

   "browser_specific_settings": {
      "gecko": {
        "id": "zhongwen@example.org"
      }
    },

 Zip the add-on into an 'xpi' file:

    $ zip -r -FS ../zhongwen.xpi * --exclude '*.git*'

Then install the add-on from the about:addons page.

 

But it took me an unreasonably long time to learn to do this because Mozilla doesn't document it very well.

Mozilla has various guides to developing add-ons but they are all oriented towards having them signed by Mozilla. They say that it is possible to install unsigned add-ons to select versions of Firefox but give only hints about what is required.

Some add-ons that install successfully as a temporary add-on, via about:debugging, cannot be installed permanently as an unsigned add-on, with the unhelpful message:

Installation aborted because the add-on appears to be corrupt.

They could have omitted "because..." - it would have made the message no less informative.

Mozilla support gets reports of the message but offers no explanation. Multiple reports, with all sorts of complex details but no overview of how an add-on might be corrupt or how to diagnose the problem, just specific trial and error advice. Mozilla is getting to be as bad as Microsoft.

The browser console has more information:

1645324229834 addons.xpi WARN Invalid XPI: Error: Cannot find id for addon /home/ian/dev/zhongwen.xpi(resource://gre/modules/addons/XPIInstall.jsm:1531:19) JS Stack trace: loadManifest@XPIInstall.jsm:1531:19

Why isn't the error message to the user: Installation aborted because the add-on does not have an id? Or some such? None the less, if one jumps through enough esoteric hoops, the information is available.

So, how to provide an ID? 

This page gives some hints. 

The workshop gives some hints. Mostly it tells you that you don't have to set an ID explicitly and when you do have to set and ID explicitly. At the very end (did you read all the way through the irrelevant details to the last sentence?) it says:

See browser_specific_settings in manifest.json for the syntax of setting the extension ID.

Could they have made it any less obvious? I don't think so. To make it this obscure, one would have to be deliberately trying to make it difficult to succeed in any way other than having the add-on signed by Mozilla, and one would have to develop and refine the obscurity with diligence.

In any case, the documentation of browser_specific_settings describes the id, including the two supported formats.

So, I edited manifest.json of my plugin, adding:

   "browser_specific_settings": {
      "gecko": {
        "id": "zhongwen@example.org"
      }
    },

They don't say anything more about this format than: a string formatted like an email address, and the guidance that if it is a real email address it will attract spam. This suggests that it doesn't have to be a real email address: it just has to have the format of an email address. There is nothing about whether or how the address is used or validated, uniqueness constraints or anything else. Really, it just describes the syntax and that is all.

I then zipped the contents of the add-on directory, excluding the .git directory, to an xpi file (which is just a zip file with a unconventional extension):

    $ zip -r -FS ../zhongwen.xpi * --exclude '*.git*'

I also set 'xpinstall.signatures.required' to false via about:config.

I was then able to install the add-on.

Why do they need an ID when they don't need it to install temporarily? Obviously they don't need it, otherwise they would need it to install the add-on temporarily. It is an arbitrary restriction, making installation more difficult without adding any obvious value and possibly without adding any value at all. One more brick in the wall.

I have installed the add-on to Librewolf and a Nightly build of Firefox. It seems to work OK.

Zhongwen

Zhongwen is an add-on for Chrome and Firefox that looks up Chinese characters. I use it on Firefox.

I was using Calibre to view e-books but then I wanted to view them in a browser so I could use the Zhongwen add-on to look up the characters I don't know.

I installed Calibre-Web and began studying some of my Chinese books that were in pdf format. It worked great. I could read along and easily look up the characters I didn't know.

But then I tried reading a Chinese book in epub format and I couldn't look up the characters. The Zhongwen add-on wasn't working, except on the title line at the top of the page.

It turns out that the Zhongwen add-on doesn't work on content in an iframe and Calibre-Web presents the contents of epub books in an iframe.

I developed an enhancement of Zhongwen (see the issue85 branch) that works on content in iframes. I have only tried it with Calibre-Web but it works fine there. With this, I can view epub books via Calibre-Web and look up the Chinese characters I don't know with the modified Zhongwen add-on.

The only remaining problem is that Firefox no longer lets me install a plugin that isn't signed by Mozilla, unless I use an unbranded release. Security is good and it would be good if requiring signed add-ons was default, but removing all possibility of installing my own add-ons is going too far. Not every Firefox user is incompetent to create their own add-on and manage add-ons safely. If someone has access to my system to manipulate my add-ons, they can just replace the Firefox executable. Requiring signed add-ons is a limitation without benefit.

It is possible to install an unsigned add-on temporarily from about:debugging. But then the add-on must be re-installed every time Firefox is restarted.  

I installed an unbranded nightly build of Firefox but still there are hoops to jump in order to install the add-on permanently. It works fine installed temporarily but there is no easy to find and follow documentation about how to install a locally developed add-on permanently, only very lengthy procedures about how to register with Mozilla and publish an add-on, which is a very involved and time consuming process to learn and exercise.

So, at the moment I am stuck with loading the add-on temporarily. It seems that Mozilla, despite a good start, has decided that making the life of users difficult and less productive is the way to go.

I have submitted a pull request to Zhongwen but thus far there has been no response from the author. Who knows when or if an update of the add-on will allow it to work on content in an iframe. I will try to be patient. I have something that works well enough for my study in the meantime, though it is a nuisance having to use a different browser.

If the add-on isn't enhanced for too long, I will consider releasing my own version, but that requires getting involved with Mozilla's publishing process, which I would rather not do. It is really unfortunate that I can't put something up on GitHub or the like that people can download, review and install for themselves.

Maybe it is time to switch to a more permissive browser. Or maybe all the major browsers are now in the business of building walled gardens. It's a sad world. I wonder how long it will be before Firefox starts making it difficult to browser sites other than Mozilla's own sites and sites of those who fund it? They have already made it impossible to browse sites for which Mozilla deems the security to be inadequate. This was good when I could review the problem and manually override where appropriate, but more recent releases give no option to proceed to the site: it is simply and utterly blocked with an unhelpful message. Attempting to browse other sites results in an alert but, thus far, one can still proceed to the site: it is just a matter of dealing with the alerts and clicking through. Add-ons? There are no practical options: only Mozilla approved add-ons are allowed. The walls are going up and, no doubt, will become higher with time, then the excuse of lack of resources will be used to bring them in to reduce the scope of support. 

It's a sad world we live in. Mozilla used to develop good, free software. It still is, to the extent that I could fork Firefox and remove the restrictions, but I don't have time for that.

Friday, January 14, 2022

gmusicbrowser on Debian bullseye

gmusicbrowser is not longer available as a package to install to Debian Linux. It was removed in 2019, due to lack of maintenance and Debian bug 912882. The root cause was that gmusicbrowser depended on libgtk2-perl, which was being dropped.

gmusicbrowser issue 57 tracks progress to migrate to gtk3 since 2013. The 1.1.99.1 release is the first based on gtk3 but there has been a year of development since then, as yet unreleased.

Fortunately, it is easy to install the gtk3 based gmusicbrowser on Debian bullseye with xfce4 desktop. It is probably as easy to install with any other desktops, but I haven't tried any of them.

$ git clone https://github.com/squentin/gmusicbrowser.git
$ cd gmusicbrowser
$ sudo make install

I didn't have to install any packages beyond what were already installed.

I haven't tested it extensively, but I haven't had any problems. If you too like gmusicbrowser, you can install it to current Debian system easily this way, and probably many other systems.

Saturday, January 1, 2022

Preparing to configure Windows

 I bought a new laptop - a dynabook tecra. It came pre-installed with Windows 10. Completing the Windows 10 setup was easy but for the next few days, every time I reboot it installs more updates.

This time, it has been displaying 'Preparing to configure Windows' for over 20 minutes with no indication of progress or when it might finish. 

The Windows Club has a page on this. They recommend waiting for 2 hours before giving up and trying anything other than waiting.

TWO HOURS! For an update. Not even a full install.

I can complete an update of Linux in about 2 minutes, typically.

What is Microsoft thinking, to make such an obtuse update process? It's not like they are beginners, without experience. They have been doing this for years.

It is because of nonsense like this that I prefer Linux. I never have such problems with Linux.

I would just abort the Windows update but, being conservative, I want to backup the Windows partitions before I wipe them and install Linux and I don't want the backups corrupted by an incomplete update. I am having to reboot a few times to confirm I have all the Linux issues sorted (modules and firmware for all the essential hardware). It is extremely annoying having to deal with Windows again.

Eventually it had completed its update and now I am unable to access the BIOS setup or boot menu. It boots straight to Windows every time, restart or shutdown and boot. Fucking Windows update - it was working normally until the fucking update.


Saturday, December 11, 2021

srf - spaced repetition flashcards

I have been using srf to study Mandarin for about 6 months now, with a selection of decks imported from anki, and cards I have created myself. 

My study time is better regulated and I am making better progress than when I was studying with anki. 

The difference is the scheduling algorithm. The scheduler in srf regulates new cards automatically, to maintain a target study time per day and sorts due cards differently, allowing good progress even with a large backlog of due cards (e.g. after not studying for a few days).

Initial development of srf is now complete. Rate of change is slowed. The user interface, particularly for editing content, remains a bit crude but adequate. Lately, I spend most of my time studying and very little developing the program.

I was reluctant to develop a new program. It took a lot of time away from study. But the limitations of the anki scheduler were too frustrating and it is now clear that a better scheduler makes a big difference to progress. From a technical perspective, the differences are subtle and simple but practically they make a big difference.

The essential differences are:

  • present cards with shorter interval before cards with longer interval
  • automatically regulate new cards based on past and future workload

Presenting cards with shorter intervals first makes a big difference when working through a backlog. Anki presents cards in the order they were due, regardless of interval. The difference is subtle but important.

With a large backlog in Anki, the backlog effectively blocks review of cards with shorter intervals. Unfamiliar / difficult cards are not seen as soon as they should be, defeating learning. With a large enough backlog, it is practically impossible to learn: one churns through a large number of cards without making progress because it is too long between reviews.

The problem is exacerbated by anki introducing more new cards every day, despite the backlog. It is possible to change the number of new cards per day manually but paying attention to this takes attention away from studying.

The scheduler in srf prioritizes the cards you are learning (with shorter intervals) over the backlogged cards. Thus, one can learn despite the backlog and that learning is the way to clear the backlog. And while there is an excessive backlog, srf does not introduce new cards, so gradually, reliably, the backlog is cleared. When it is cleared, new cards are introduced again.

With these changes, study time per day is more consistent: varying closely around configured target study time per day. And progress is better when one is not overloaded.

libinput - touchpad acceleration - piecewise linear profile

I have implemented a piecewise linear acceleration profile for touchpad in a fork of libinput.

The mapping function is defined by an array of points: (speed, factor), which must be sorted in order of increasing speed. Between the points, factor is determined by linear interpolation.

A single point gives a fixed acceleration factor, like the libinput flat profile.

Two points give a single, linear slope, clipped at the upper and lower input speeds. But the clipping is irrelevant if the speeds of the two points are sufficiently low (e.g. 0) and high respectively, in which case all inputs will be at speeds between the points.

Factor
^
|                   _________
|                 /
|                /
|               /
|-----------/
|
+-----------------------------------------------> Speed

More complex curves can be approximated by more points. With enough points, any of the profiles in the old X.org server can be approximated.

At the moment, configuration parameters are hard coded, because I haven't figured out how to add configuration parameters and because using new parameters would require changing code outside libinput (e.g. the X or Wayland libinput drivers).

One of the limitations of libinput is that the only means of configuration is via the API. As a result, to add new configuration parameters requires modification of both the libinput library and whatever software (e.g. Wayland compositor, xf86-input-libinput, etc.) calls it. I may add configuration via environment variable to work around this limitation but in the meantime, as you will have to compile and install this version anyway, it is almost as easy to edit the source as a separate configuration file or environment variable.

See the touchpad-pl branch of this fork of libinput.

To try this profile for your touchpad:

$ git clone https://github.com/ig3/libinput.git
$ cd libinput
$ git checkout touchpad-pl
$ meson --prefix=/usr builddir/
$ ninja -C builddir/
$ sudu ninja -C builddir/ install

This is working for me on Debian 11 with xfce desktop. I haven't tried a system running Wayland.

To change the profile, edit src/filter-touchpad-pl.c. The parameters are in function touchpad_accel_profile().

To test:

$ sudu libinput debug-events

Labels