Systems · 2026
Setu — a club's WhatsApp presence that answers from the record
SANKALP's assistant on WhatsApp. It answers questions from the constitution and the filed record without a language model, carries out a fixed set of acts on the club's platform on a member's word, holds everything else for an office holder to approve, and never holds a production credential.
SANKALP is a student organisation with a constitution, a filed register of its decisions, and a platform I built for it. Members ask the same questions in WhatsApp groups over and over — how long is a term, who can call a meeting, what does the charter say about this — and the answers are all written down somewhere nobody looks. Setu (bridge) sits on a club number in the executives' group and in direct messages and answers from that record. Then, because a bot that can only quote is a bot nobody uses, it learned to act.
Everything is recorded, little goes further
Every message in every group it is in is recorded first, unconditionally, into a local capture. A router in ordinary code then decides what goes further: only a message that names Setu, tags it or replies to it, and only in the executives' group or a direct message. Chatter of five words or fewer is dropped before identity or the production platform is touched.
A question is answered without a model at all: a curated map of about twelve hundred questions, BM25 retrieval over passages rather than whole documents, the platform's own rules deciding per asker what may be quoted, and verbatim text with a citation. An instruction is classified by whole phrases into one of five acts the charter allows, and only then does a Claude Code session start — with the act, the actor and the one authorised script already fixed, and the member's words quoted as untrusted material. Each script's hash is its authorisation; edit one by a character and it needs a human again. Anything the classifier cannot name needs an office holder to reply approved to the refusal, in the thread.
Three things it never has
It never holds a production credential: it reads an hourly mirror of the platform's database as a user with SELECT on one schema and nothing else, so the scoping is the credential rather than a promise about what the code happens not to call. It never opens a thread: a reminder lands in the one group it was told to speak in or a DM the member opened, and a member with no DM open is simply not reachable. And it never touches my own databases, which are not configured in its process at all.
The first afternoon was wrong five times out of five
Nothing failed. Asked when the last meeting of a wing was, it quoted the article defining the wing. Asked to create a task, it quoted a scanned report. Told "I love this guy", it offered to continue in a DM. Two causes: nothing was allowed to say this is not a question, and a word the map had never seen weighed nothing, so one common word alone could carry a 62% match.
The fix was all code, and it was measured against a held-out set written afterwards — 45 questions with the article that answers them, 22 pieces of ordinary talk taken from the group, 15 routings. Right article in the top three went from 30 to 38 of 45; talk answered with silence from 0 to 21 of 22; routed as a member would expect from 5 to 15 of 15. The set is committed so the number can be re-run.
A local model, kept on a short leash
A small model runs locally on the machine's own GPU, and it is heard only where code was not sure: an act verb with no matching phrase, a message with no shape, a question whose kind the patterns missed. It answers with a label out of a JSON schema and nothing else. It never lifts a refusal, never writes a word a member reads, and never sees captured text as anything but a quoted block. It was chosen by measurement across three candidates on thirty-six held-out messages, and the prompt improved most when it was asked to copy each detail out of a message rather than judge whether the detail was there — code decides presence.
For questions the record cannot answer, a session with one tool — the platform's own dashboard, as the asker, with no way to name anyone else — looks properly and says so first. And an office holder can ask for the full thing by name, which goes past the router and the map to a session bounded not by its prompt but by its four tools, so the platform's own permission checks decide what that member may do. Probed by a session asked to break out: seven refusals, three reads.