# blog.jaye.ch $ grep -r▊

The search

every published piece, by three signals at once — the words, the vectors, and the links between them.
# 61 pieces · 16168 words in the list · 239 citations from 184 works · three signals: the words, the vectors, the links
  1. Type a word — the pieces are searched by their words, their vectors and their links.

This page searches all 61 published pieces at once and tells you which of three signals put each result in front of you. The first is the words: the terms the pieces actually use, three characters and longer, each with the pieces it occurs in — looked up as a whole word and as the beginning of one, where a hit in a piece's title or a section heading counts for more than a hit in its body. The commonest words are left out (they are in every piece, so they name none); everything else is here, including a name used by a single piece and only once in it.

The same list takes a pattern: put one between slashes and it is matched against the words themselves rather than looked up — /wetiko.*/, /^procl/, /(gno|the)sis/, /urgy$/. With nothing anchoring it, a pattern matches anywhere inside a word, which is how a half-remembered word is found: /theurg/ reaches it wherever it sits. The syntax is . for any character, * + and ? for repetition, [abc], [a-z] and [^abc] for a set of characters, | to alternate between expressions, () to group them, and ^ and $ for the start and the end of a word. A pattern that cannot be read is refused in the line above the results rather than guessed at, and a pattern reaching more words than the page will match is cut off there and says so.

Put a ~ in front of a word and it is matched approximately instead of looked up: ~proculs finds proclus, because one tilde allows one edit — a wrong letter, a letter too many or too few, or two letters the wrong way round, which is what most typos are. ~~proculs allows two edits, and reaches several times as many words for it. A word found this way is not a word you asked for, so it always ranks below a word you did type, and the reason under the result says which word it was and how far away.

The box also finishes words as you type them: three letters or more and the words of the list that begin that way are offered beneath it, up to eight, and the list says so when it stopped there. The arrow keys move through them and Enter takes the one you are on; Escape puts the list away and leaves what you typed. When a query finds almost nothing and one of its words sits one edit from a word the list holds, the line above the results offers that word with the misspelling left in place — the page suggests, it never rewrites. The results take the same arrow keys once no list is open, Enter opens the piece the selection is on, and Home and End go to the first and the last. A word the query reached is marked in the result's title as well as in its snippet, and the order can be switched from relevance to newest first — the mode travels in the link, so a sorted search is one you can send.

Words between quotes are a phrase: "alpha beta" matches only the pieces where those two words stand NEXT to each other, in that order — every quoted run is its own demand, and the loose words outside the quotes stay OR-ed as always. A one-word quote is that word; a phrase the pieces never say that way matches nothing, and the line above the results says which phrase came to nothing rather than padding the list with its words' separate hits. Opening a result's passage and marking it where the match sat needs the pieces' own text, which is fetched once from this site on the first search and kept for the visit — until it arrives the page searches exactly as it does without it, from the headings, summary and opening, and a fetch that fails changes nothing else.

The second is a vector space — 160 dimensions per piece, built by hashing each term to a dimension and a sign. It is a tf-idf projection and not a neural embedding: there is no model here, and no projector either — this page recomputes your query's projection with the same hash and takes the cosine. That is how a piece that argues in different words still surfaces. Only words the pieces really use are projected, because a word none of them uses would land in dimensions the pieces' own terms have already filled and would answer with noise; a query of such words returns nothing, and the page says so rather than padding the list.

The third is the blog's own graph. Every piece is a node, every link one piece makes to another is an edge, and the 184 works the pieces cite — the books, papers and primary texts named in their notes — join pieces that never mention each other. Those facts are read by a small datalog evaluator running in this page, with its rules declared as data: the reasons under each result are the rules that fired, and a result can be here on the strength of the graph alone. The expansion is capped (the strongest few hits seed it and the derived facts are bounded), so an unusual query cannot make the page wait.

Nothing is sent anywhere: the index is in this page and the search runs in your browser, with the passages fetched once from this site. Snippets come from each piece's own words: the passage a match sat in when the fetched half is there, and each piece's headings, summary and opening passage when it is not — so a piece matched deep inside shows its opening rather than the matched line. Scores are relative to the query and are not a judgement of the piece.