A lightweight Express service that schedules and sends Firebase Cloud Messaging (FCM) & text (SMS) push notifications to wake up registered devices.
Device registrations are stored in a local SQLite database.
Quick Start
cp .env.example .env
Edit .env — set FIREBASE_SERVICE_ACCOUNT_JSON.
Here is one way to generate the contents: cat your-downloaded-key.json | jq -c .
Optionally set ENDORSER_URL / PARTNER_URL if you are not using the production Endorser (https://api.endorser.ch) and Partner (https://partner-api.endorser.ch) hosts.
Optionally set NOTIFY_DATA_DIR if you want the SQLite database somewhere other than ./data.
pnpm install
pnpm run dev
pnpm test
The server starts on http://localhost:3003 (or the port in PORT). Hot-reloads on file changes. pnpm start is the same tsx entry without watch. These commands are not the production Docker path (node dist/index.js).
On first use, the service creates NOTIFY_DATA_DIR (default ./data) and the SQLite file notify.sqlite with the required schema.
FCM
Authentication
POST /notifications/register and POST /notifications/refresh require a Bearer JWT. After local JWT verification, the service checks the token with Endorser. Registration and refresh continue only if Endorser accepts the JWT.
PUT /notifications/alert-authorization uses the same current-user Bearer JWT + Endorser check. It does not accept the testMode local bypass. The 100 delegated JWTs in the body are stored credentials, not the request authenticator.
Local notification test bypass: send testMode: true in the JSON body and omit the Authorization header. The request skips JWT and Endorser checks and uses a synthetic local test user, same as before. This applies to register/refresh only.
Set NODE_ENV=test-local in .env to bypass ethr JWT expiry verification during local development (this is separate from the testMode bypass above).
Alert authorization
PUT /notifications/alert-authorization
Authorization: Bearer <current-user-JWT>
{
"batchId": "client-batch-id",
"notifyHourUtc": 18,
"notifyMinuteUtc": 30,
"timezone": "America/Denver",
"jwts": [
{
"sequence": 0,
"day": "2026-08-27",
"nbf": 1756252800,
"exp": 1756339200,
"jwt": "eyJ..."
}
]
}
day is a UTC calendar day, and each JWT must be valid for the whole of the day it names: nbf at or before midnight UTC that opens it, exp at or after midnight UTC that closes it. The daily run selects by UTC day and may fire at any moment inside it, catch-up runs included, so a window covering only part of that day would hand Endorser a credential outside its own validity period. Days must be distinct and the 100 sequence values consecutive.
notifyHourUtc and notifyMinuteUtc are required and timezone is optional; all three behave exactly as on the SMS twin. Both channels gate on the stored hour (Notification hour).
A successful call replaces that user's previous unused JWTs atomically. Passkey (did:peer) identities cannot mint this batch and receive DELEGATED_JWT_UNSUPPORTED_IDENTITY.
Removing an authorization
DELETE /notifications/alert-authorization
Removes every batch and every JWT this DID holds for the push channel, consumed rows included, so the alertSearch scheduler stops listing the identity. It answers { success: true, deletedBatches, deletedJwts }, with zeros when there was nothing stored. Device registrations and WAKEUP_PING are untouched, and so are the alertSearch cursors: a later re-authorization resumes where this one stopped rather than replaying months of history.
Alert search retrieval
The daily scheduler runs retrieveAlertSearch against:
{ENDORSER_URL}/api/v2/report/alertSearch{PARTNER_URL}/api/partner/alertSearch
The delegated JWT is sent as Authorization: Bearer. Pass independent endorserAfterId / partnerAfterDate (or omit them on first run). Nearby search uses the alertSearch bbox (minLocLat, maxLocLat, minLocLon, maxLocLon).
loadAlertSearchCursors / retrieveAlertSearch / advanceAlertSearchCursors (or runAlertSearchCycle) persist those bounds per user DID in SQLite. Cursors advance only after a complete success retrieval (not empty, pagination, or errors). A Partner page of 50 rows that share the oldest updatedAt is pagination because exclusive beforeDate cannot drain timestamp ties.
runDailyAlertSearch(userId, now?) picks the unused delegated JWT for the current UTC day, runs runAlertSearchCycle with that JWT, and marks that specific JWT consumed only when both Endorser and Partner complete (success or empty, including both empty). Pagination or source failures leave the JWT unused so the same day can be retried. The result reports the day it used as utcDay. After a retrieve, the result includes digest from buildAlertSearchDigest (six bucket records and counts). digest is null when there is no batch or no unused JWT for today. Consumption does not depend on digest.hasUpdates.
startAlertSearchScheduler() (started from src/index.ts next to the FCM scheduler) is a separate user-level job. One query per pass asks for the users who hold an unused JWT for the current UTC day, flagged by whether their notify_hour_min_utc has arrived (Scheduler selection); it calls runDailyAlertSearch on the due ones, up to ALERT_SEARCH_USER_CONCURRENCY at a time. After each run, if the digest is complete with updates and today's JWT was consumed, it sends a user-visible FCM message (title: TimeSafari, body You have N new updates., data type: alert_search) to that user's registered tokens. It does not call sendPushToDevice or change WAKEUP_PING. Subsequent ticks the same UTC day see no unused JWT (digest: null) and do not resend. FCM send failures are logged and do not roll back cursors or JWT consumption. A process-local in-flight flag skips a tick if a pass is still running.
buildAlertSearchDigest maps a retrieve result into structured payload data: per-bucket record arrays, counts, totalCount, hasUpdates, and Endorser/Partner completion status. It does not invent a notification string; FCM uses only totalCount for the short body. Incomplete outcomes (pagination, auth, network, etc.) yield completed: false and hasUpdates: false. Empty successful retrieves are complete with hasUpdates: false and do not send FCM.
SMS
/notify-sms delivers the daily alertSearch digest by text as well as by push.
A user registers a phone number, proves possession of it with a 6-digit code,
authorizes a batch of delegated alertSearch JWTs for the SMS channel, and
receives at most one text per UTC day when that day's retrieval finds updates.
The whole surface is off unless SMS_ENABLED is true; every route answers
503 SMS_DISABLED otherwise, and the SMS scheduler does not start.
Endpoints
| Method | Path | Purpose |
|---|---|---|
| GET | /notify-sms/phone |
List this DID's registrations; with ?phoneNumber=, the other DIDs on a number this DID has verified |
| POST | /notify-sms/phone |
Record a phone for the DID, unverified, and text it a 6-digit code |
| PUT | /notify-sms/phone |
Match the code and mark the registration verified |
| DELETE | /notify-sms/phone |
Remove the phone entirely |
| POST, PUT | /notify-sms/alert-authorization |
Store a delegated JWT batch for the SMS channel |
| DELETE | /notify-sms/alert-authorization |
Remove every SMS batch and JWT for the DID, turning the channel off |
| POST | /notify-sms/inbound |
Twilio's webhook for STOP / START / HELP |
GET /notify-sms/phone returns the caller's own numbers in full. Adding
?phoneNumber= also returns dids, every DID holding a verified registration
of that number — but only to a caller who has itself verified that number.
Otherwise 403 SMS_PHONE_NOT_VERIFIED_BY_CALLER, with no count and no
identities.
POST /notify-sms/phone takes { "phoneNumber": "+15555550123" }. A number the
DID has already verified returns 200 with verified: true and sends nothing.
Otherwise it mints a code, stores the code's HMAC, and texts the plaintext. The
response carries { success, phoneNumber: "+1555*****23", verified: false, expiresAt } and never the code. Sends are throttled to three per phone per hour
(counted across every DID) and ten per DID per day.
PUT /notify-sms/phone takes { "phoneNumber", "code" }. A miss returns
400 SMS_CODE_MISMATCH with attemptsRemaining; SMS_CODE_MAX_ATTEMPTS misses
return 429 SMS_CODE_ATTEMPTS_EXHAUSTED and clear the code, so recovery is
another POST.
DELETE /notify-sms/phone accepts the number in the body or as ?phoneNumber=,
because a fair number of proxies drop bodies on DELETE. It removes the row
matching (user_id, phone_e164) and sets phone_e164 to null on that DID's
sms_phone_log rows for that number, leaving phone_hash and the action
history intact. Deleting a number that is not registered returns
{ success: true, deleted: false }, not an error.
POST /notify-sms/alert-authorization takes the FCM twin's body plus a
required notifyHourUtc (integer 0-23) and notifyMinuteUtc (integer 0-59):
the hour the user wants to hear from the service. Both are UTC, which the
field names carry, so no offset and no zone travel with them and there is
nothing for the server to misread. A client sends date.getUTCHours() and
date.getUTCMinutes() with no formatting step.
Both are required, and one without the other is refused naming the missing
field, so every batch states its own hour rather than inheriting one
(Notification hour). Out of range, fractional, or the
wrong type is rejected with the rest of the batch. The response echoes both
integers; the server composes the zero-padded HH:MM it stores.
An optional timezone (an IANA name such as "America/Denver") is validated
and stored beside them. Nothing reads it: the two UTC integers are what
schedule a send. It is recorded against the day something needs to re-derive
that hour across a DST change — see Notification
hour. A batch may carry the hour without it, but not the
reverse. A name Intl cannot resolve is rejected with the rest of the
batch, because a zone this service cannot resolve would be worth nothing to
whatever reads the column later.
The route requires at least one verified phone for the DID;
without one it returns 409 SMS_NO_VERIFIED_PHONE. PUT is accepted as an
alias, since the FCM twin is PUT and the semantics are replace-not-append
either way.
DELETE /notify-sms/alert-authorization is how a user turns texts off. It
removes every SMS batch and JWT for the DID, used or not, so the SMS scheduler
stops listing that identity, and answers { success: true, deletedBatches, deletedJwts } — zeros when there was nothing stored. Registered phone numbers
survive: silencing alerts is not a request to redo the possession check later.
DELETE /notify-sms/phone is the route that forgets a number, and STOP is the
route that blocks one. The revocation is recorded in sms_phone_log as
alert-authorization-deleted; for a DID whose handset is already gone, the
phone_hash on that row stands in for the number the column expects.
Action claim
Every /notify-sms call is authorized by a Bearer JWT whose claim names the
action and the phone number it applies to:
{
"iss": "did:ethr:0x…",
"iat": 1756270800,
"exp": 1756271100,
"claim": {
"@context": "https://giftopia.tech",
"@type": "SmsNotificationAction",
"action": "register-phone",
"phoneNumber": "+15555550123"
}
}
action is one of list-phones, register-phone, verify-phone,
delete-phone, authorize-alert-search, revoke-alert-search. phoneNumber
is required for register-phone, verify-phone and delete-phone, and
required for list-phones only when the request carries the query parameter.
authorize-alert-search and revoke-alert-search act on the DID's whole
inventory rather than on one handset, so they bind to no number. The claim holds
no DID — the authenticated identity is iss — and no batchId.
@context is https://giftopia.tech with no trailing path, shared with the FCM
setup claim; @type is what separates them. https://giftopia.me is the app
link that appears in messages and is never the claim namespace.
The middleware chain is requireAuth (Bearer JWT, signature verified against
the issuer DID), requireEndorserAuth (Endorser accepts the same token), then
requireSmsActionJwt(action), then the handler. The handler performs no
authorization checks of its own. There is no testMode bypass: no local-test
path sends real texts to real handsets.
| Failure | Response |
|---|---|
| No claim in the token | 403 SMS_ACTION_JWT_MISSING_CLAIM |
| Claim names a different action | 403 SMS_ACTION_JWT_WRONG_ACTION |
| Claim names a different number | 403 SMS_ACTION_JWT_PHONE_MISMATCH |
iat outside SMS_ACTION_JWT_MAX_AGE_SEC |
401 SMS_ACTION_JWT_STALE |
exp has passed |
401 SMS_ACTION_JWT_EXPIRED |
| Token already used | 401 SMS_ACTION_JWT_REPLAYED |
The token's sha256 is recorded before the handler runs, so one token buys one
action. A handler that fails afterward does not release the hash; the client
mints a fresh JWT, which it can do freely. SMS_REQUIRE_ACTION_CLAIM=false
turns the stage off entirely.
This is a client change. The TimeSafari app sends a plain identity JWT to
/notifications/*; a /notify-sms call carrying no claim returns
SMS_ACTION_JWT_MISSING_CLAIM.
One phone, several DIDs
Every read and write is scoped by (user_id, phone_e164), so two identities
sharing one handset stay independent: registering, verifying, or deleting under
one DID does not touch the other's row, and each identity runs its own
alertSearch against its own cursor. Two verified DIDs on one handset therefore
receive two texts a day.
SMS_MAX_DIDS_PER_PHONE bounds how many DIDs one handset can carry, counting
verified rows only — counting every row would let five throwaway DIDs lock
the handset's owner out by registering and never verifying. The count is checked
at POST as an early rejection and again at PUT, which is the check that holds,
since PUT is the moment a row starts consuming a slot.
A blocked POST returns 409 SMS_PHONE_DID_LIMIT with limit and
verifiedCount and no identities: a POST names any phone number on earth
and proves nothing about it, so answering with DIDs would make the endpoint a
phone-number-to-identity lookup oracle. A blocked PUT returns the same code plus
dids, because a PUT that reaches the limit check has already matched a correct
code and the caller is holding the handset. The code is consumed on a limit
rejection exactly as on success, so a fresh answer costs a fresh POST.
Carrier opt-out is the one place the DID boundary is crossed on purpose: STOP
arrives with a phone number and no identity, so it marks every registration of
that number unverified. START does not restore anything; possession has to be
proved again with a fresh POST and code.
Blocking a number
sms_blocked_numbers is the service's suppression list. A number on it is
refused at registration (403 SMS_PHONE_BLOCKED), refused at verification even
with a correct code, and skipped by the daily digest. Deleting a registration
still works — removing yourself is always allowed.
Who blocked a number
The reason column records the origin, and nothing ever downgrades it:
reason |
Origin |
|---|---|
opt-out |
The handset texted STOP to this service |
provider-opt-out |
Twilio refused a send with 21610, so the number is on its own opt-out list — it opted out somewhere this service did not observe |
manual |
An operator added it |
A re-block keeps whichever reason carries the stronger claim
(opt-out > provider-opt-out > manual). An operator re-blocking a number
that had already opted out must not erase the opt-out, because that record is
what says the block may not simply be lifted again.
pnpm run sms:blocks lists the table grouped by origin, sms:blocks list manual
filters to one, and sms:blocks block <number> "<why>" / sms:blocks unblock <number> manage entries without hand-written SQL — which also keeps reason
honest, since a manual INSERT can claim any origin it likes. Unblocking
anything other than a manual entry requires --force.
SMS_CODE_SECRET=... pnpm run sms:blocks list
Two sources are consulted. The table is written by the STOP webhook, and
SMS_BLOCKED_NUMBERS in the environment blocks a number by configuration without waiting for the
handset to ask:
SMS_BLOCKED_NUMBERS=+15555550123,+15555550124
Entries are normalized before comparison, so formatting cannot defeat them.
Rows are keyed on phone_hash, so a block outlives the DELETE that nulls
phone_e164 elsewhere. This table does retain the number itself: a suppression
list nobody can read is one nobody can audit or correct, and honoring an
opt-out means remembering who asked.
STOP blocks as well as unverifying. Unverifying alone left the number free to
register again minutes later and resume queueing messages, which Twilio then
refuses with 21610 — the service would retry forever on behalf of someone who
asked to be left alone. START lifts the block, matching Twilio, which clears
its own opt-out entry on the same keyword. It does not restore verification:
possession was proved by a code and that code is gone, so a fresh POST and a
fresh code are the way back.
Twilio keeps its own opt-out list regardless, and it is the authority for
compliance. A send refused with 21610 is Twilio saying the number is on it, so
both send paths record that as a provider-opt-out block and stop retrying.
Twilio can also report 21610 asynchronously; catching those needs the
StatusCallback handling this service does not yet do, so synchronous refusals
are what get captured today.
An ordinary send failure never blocks — only 21610 does, so a carrier blip
stays retryable.
Restricting who a server may text
SMS_ALLOWED_RECIPIENT_DIDS is a comma-separated list of the only DIDs an
instance may send to. Unset — the production case — there is no restriction.
SMS_ALLOWED_RECIPIENT_DIDS=did:ethr:0xabc,did:ethr:0xdef
It guards both paths that spend money: the verification code on
POST /notify-sms/phone, which returns 403 SMS_RECIPIENT_NOT_ALLOWED and
registers nothing, and the daily digest, which is withheld before any send.
Either way an recipient-not-allowed row lands in sms_phone_log, one per
number, so the log records exactly what was held back. Matching is
case-insensitive, since a checksummed did:ethr address and its lowercase form
name one identity.
Set but empty blocks every recipient rather than allowing all of them. An operator who sets the variable meant to restrict something, so the blank case fails closed.
This exists because the dangerous configuration is a test server holding live
Twilio credentials and a copy of the production database: every
sms_registrations row with verified = 1 is a real handset, and the scheduler
texts all of them daily. The allowlist turns that from an incident into a log
line. SMS_ENABLED=false remains the blunter switch — it returns 503 from
every route and stops the SMS scheduler from starting at all.
Scheduler selection
Both alertSearch passes choose their users with one query, not one query per
user. listPendingForDay({ day, hourMinute }) returns every user holding an
unused JWT for that UTC day, each flagged due by whether their batch's
notify_hour_min_utc has arrived:
- A user who has already run today holds no unused JWT for it and does not appear at all, so the ~287 ticks a day that have nothing to do for them cost nothing.
- A user whose hour has not arrived appears with
due: falseand is counted indeferred, which is what keeps that number in the log line. Every batch has an hour, so this applies to every user rather than to a subset. - The batch consulted is the newest one per user, picked with a window function. A user accumulates batch rows, because a batch with a consumed JWT survives the next upload, so "the user's notify time" is not a plain join.
HH:MMis zero-padded, so the text comparison SQLite performs is chronological.
Due users are then worked on ALERT_SEARCH_USER_CONCURRENCY at a time
(src/util/concurrency.ts, currently 8). The per-user work is two external
round trips against that user's own cursors, so a serial loop spends the pass
waiting: at 200ms per user it stops fitting inside a five-minute tick at roughly
1,500 users, after which the in-flight guard skips passes and users miss days.
The bound is held low deliberately — the ceiling it relieves is latency, and the
two APIs on the other end are shared infrastructure that a wide fan-out would
only move the queue into.
Measured on this codebase at 2000 users, one tick: a pass where every user has already run costs 0 queries and ~1ms (it was 3 queries per user and ~141ms when the scheduler asked per user), and a pass where every user is due drops from ~106s to ~13s at 50ms of API latency.
SMS delivery
startSmsAlertSearchScheduler() is a third scheduler alongside the FCM wakeup
and FCM alertSearch passes, with its own interval
(SMS_ALERT_SEARCH_INTERVAL_MS, default 5 minutes) and a 150-second initial
offset so the two alertSearch passes do not hit Endorser in the same instant. It
selects its users with the one query described in Scheduler
selection, calls runDailyAlertSearch(userId, now, {}, "sms") on the due ones with a bounded number in flight, and then
deliverAlertSearchSms. It logs [SmsAlertSearchScheduler] Pass started /
Pass completed in, the latter with attempted, deferred, and failed
counts. The in-flight guard is process-local, exactly like the other two, so a
second replica double-texts.
Eligibility is the same predicate the FCM path uses: the run consumed today's
JWT and the digest is complete with updates. Consumption is what makes later
ticks on the same UTC day no-ops. The message is
Gift Economies: you have N new updates. https://giftopia.me Reply STOP to end.,
kept inside one 160-character GSM-7 segment, since a second segment is a second
charge. Underneath the JWT rule, an sms_phone_log count caps sends at one per
handset per identity per UTC calendar day. Send failures are logged and do not roll back
cursor advancement or JWT consumption.
The SMS channel keeps its own JWT inventory and its own cursor table. A user on both channels produces two Endorser and two Partner queries per day and needs 200 minted JWTs; the counts can differ transiently when one channel's retrieval fails and the other's succeeds. A shared cursor was rejected: two independent daily runs against one row means whichever fires first consumes the delta and the other reports nothing.
Notification hour
Every batch on either channel carries a UTC hour and minute. The server stores
them zero-padded as one HH:MM value in notify_hour_min_utc, which is the
form SQLite compares chronologically. Both schedulers hold a
user's whole daily run until that UTC time arrives, and count the held users as
deferred in the pass result. The decision is made in SQL, not per user — see
Scheduler selection.
The field is required rather than optional because the alternative default is not "no gate" but "the first tick after midnight UTC" — a user's JWT for a new UTC day becomes selectable at that moment, so every user without an hour runs in the same five-minute window. Making each batch name its own hour spreads that load as a side effect of asking the question.
The gate sits ahead of the search rather than ahead of the text. Running the search consumes that UTC day's JWT, and a digest is sent only for the run that consumed it, so a search that ran at the top of the day would leave nothing to send at the chosen hour.
The hour lives on the batch rather than inside the delegated JWTs. Those are the
alertSearch credential, and their nbf/exp bound a whole UTC day; narrowing
them to an hour would narrow when the search may run against Endorser, not when
the user hears about it.
A stored UTC time does not follow the user through a daylight-saving change.
Someone in Denver who wants 18:00 local sends 00:30 UTC in summer, and when
their region returns to -07:00 that same UTC instant reads 17:00 on their wall
clock. The correction available today is a fresh batch carrying the new UTC
hour — which a
client uploads roughly every 100 days anyway, since that is how long an
inventory lasts.
The optional timezone on the batch exists for a mechanism that would close
that gap without waiting for the next upload, by re-deriving notify_hour_min_utc
from the zone's current rules. No such mechanism runs: the column is recorded
and unread, and the schema says so in a comment SQLite keeps, so .schema shows
the reason next to the column. Which mechanism it should be is open — a job that
sweeps changed zones, or a stored next-firing instant recomputed each time a
user fires, which needs no scheduled job at all.
There is no upper bound within the day. A service that was down at the chosen
hour and comes back six hours later still runs that UTC day; a silent day is the
worse failure. The text lands on the first tick at or after the hour, so within
one SMS_ALERT_SEARCH_INTERVAL_MS in the ordinary case. An hour late in the UTC
day leaves a correspondingly short window before the day key rolls and that
day's JWT is skipped.
A row whose notify_hour_min_utc is NULL or not HH:MM is treated as due: a value
nobody can read must not silence a channel the user asked for. The selection
query enforces that with a GLOB guard, because text ordering alone would rank
an unreadable value above every real HH:MM and defer such a user permanently
rather than once. The route cannot produce either state — only a write that
bypasses it can.
Testing SMS locally
Two scripts cover the two things worth checking separately. Both default to fake data and neither needs a real handset, a purchased number, or 10DLC registration.
pnpm run sms:send [to] [body] makes one send and prints the result. No server,
no database, no auth — just the Twilio path. The destination can also come from
SMS_SEND_TO.
TWILIO_ACCOUNT_SID=ACxxxx TWILIO_AUTH_TOKEN=xxxx \
TWILIO_MESSAGING_SERVICE_SID=MGxxxx \
pnpm run sms:send +15551234567 "test from my Mac"
pnpm run sms:smoke [to] runs the whole route flow: POST, PUT with the echoed
code, both GET forms, DELETE, then a dump of sms_phone_log. The destination
can also come from SMS_SMOKE_TO. It stubs the two things
that otherwise need the real world — it starts a throwaway Endorser that answers
/api/report/rateLimits with 200, and mints unsigned did:ethr JWTs, which
decodeAndVerifyJwt accepts under NODE_ENV=test-local without checking a
signature. The middleware chain, the claim check, the throttles and the database
are all real. Each run gets a fresh NOTIFY_DATA_DIR, so the three-codes-per-
hour throttle never interferes.
With no Twilio credentials set, sends go to the console adapter and nothing leaves the machine.
pnpm run twilio:whoami answers "whose account am I about to bill?" — it
fetches the Account resource with the configured SID and token, which separates
a mismatched credential pair from a working one before any message is involved.
It sends nothing and costs nothing.
TWILIO_ACCOUNT_SID=ACxxxx TWILIO_AUTH_TOKEN=xxxx pnpm run twilio:whoami
200 prints the account's friendly name, status and type, and that is the
account a send would bill — note type: Trial can only reach verified numbers.
401 / 20003 means the SID and token are not a matching pair. 403 / 20008
("Resource not accessible with Test Account Credentials") means the pair is a
valid test pair: test credentials may not read the Accounts resource, so
that refusal is a pass, not a fault.
Twilio test credentials are the cheapest way to exercise the real API, and they behave the same whether or not a 10DLC campaign is approved — nothing they send reaches a carrier. They are a separate Account SID and Auth Token from the live pair, under Console → API keys & tokens → Test credentials; a live SID with a live token sends real, billable messages. They are a second Account SID / Auth Token pair in the Twilio console, separate from the live ones; they need a (free) account but no purchased number, they deliver no message, they trigger no status callbacks, and they cost nothing.
TWILIO_ACCOUNT_SID=ACxxxxtest TWILIO_AUTH_TOKEN=xxxx \
TWILIO_FROM_NUMBER=+15005550006 pnpm run sms:smoke +15551234567
+15005550006 is the only From that passes validation; every other number
returns 21606.
The To is validated even under test credentials, so a reserved fictional
number such as +15555550123 is rejected with 21211. Both scripts refuse to
run with that placeholder once Twilio is configured, rather than spending a
round trip to learn it.
The destination is normalized before anything is sent, and both scripts echo
the result — check that line first when Twilio rejects a number. Ten digits with
no + are assumed US, so 8015601471, 801-560-1471 and +18015601471 all
reach the same place. A + prefix is taken at its word: +8015601471 is
syntactically valid E.164 with country code 80, so it passes normalization and
is rejected by Twilio rather than here.
A real To under test credentials is less predictable: some accounts accept
it and return a synthetic SID, others answer 20404
(resource ... Messages.json was not found) despite the credentials being
valid. Treat the magic To numbers below as the dependable path for test
credentials, and use live credentials when a text has to actually arrive.
These magic To numbers force specific failures, useful for exercising the
code-send-failed path on purpose:
To |
Twilio error |
|---|---|
+15005550001 |
21211 invalid number |
+15005550002 |
21612 cannot route |
+15005550003 |
21408 no permission for that region |
+15005550004 |
21610 blocklisted |
+15005550009 |
21614 not SMS-capable |
Even with deliberately wrong credentials the round trip is worth running once:
Twilio answers Authentication Error - invalid username, which proves the URL,
the Basic auth header, the form encoding and the response parsing all work and
only the credentials are missing.
Sending to a real handset needs a real (trial or paid) account, a real From
number, and — for a US long code — completed A2P 10DLC registration.
Provider
Sends go to Twilio over plain fetch against
https://api.twilio.com/2010-04-01/Accounts/{SID}/Messages.json with HTTP Basic
auth and a form-encoded To / From (or MessagingServiceSid) / Body. There
is no twilio SDK dependency. With configuration absent or incomplete, sends
return SMS_NOT_CONFIGURED and the process still boots: a texting outage must
not take push down with it. Under NODE_ENV=test-local with no Twilio
credentials, a console adapter prints the message instead of sending it.
SMS_DEV_ECHO_CODE adds a devCode field to the POST response holding the
plaintext six digits, so a developer with no Twilio account or no carrier
coverage can still exercise POST-then-PUT. It is honored only when
NODE_ENV is also test-local, checked first, so a production process with the
flag set by accident echoes nothing.
US A2P 10DLC registration is required before Twilio will carry
application-to-person traffic on a long code. Brand and campaign registration
take days and carry per-campaign fees. Unregistered traffic gets filtered by
carriers silently, with a sent status from the API.
A registered campaign lives on a Messaging Service, and every number in that
service's sender pool inherits the campaign — including numbers added later.
Set TWILIO_MESSAGING_SERVICE_SID rather than TWILIO_FROM_NUMBER once a
campaign is approved. Both deliver, since the pool carries the registration
either way, but a bare From leaves the Messaging Service off the message
record in Twilio's logs and Insights, and it makes it possible to point at a
number outside the pool and quietly send unregistered traffic. The Messaging
Service also picks the sender for each destination. When both variables are set
the Messaging Service wins and TWILIO_FROM_NUMBER is ignored.
A sent status means Twilio accepted the message, not that a handset received
it. This service records alert-sent on that acceptance and does not register a
StatusCallback, so delivered / undelivered / failed outcomes are not
tracked. That is a gap to close if delivery receipts matter.
The inbound webhook authenticates by Twilio's X-Twilio-Signature over the
exact URL Twilio posted to, not by JWT — it is Twilio calling, not a user. Set
TWILIO_WEBHOOK_URL to that public URL; behind a proxy or tunnel the request's
own headers do not reliably reproduce it.
Storage
Database location
| Path | Description |
|---|---|
{NOTIFY_DATA_DIR}/notify.sqlite |
Primary SQLite database (default dir: ./data) |
{NOTIFY_DATA_DIR}/notify.sqlite-wal |
WAL journal (present while the process is running) |
{NOTIFY_DATA_DIR}/notify.sqlite-shm |
Shared-memory file used with WAL mode |
NOTIFY_DATA_DIR defaults to ./data (relative to the process working directory). The data/ directory is gitignored.
Production must keep these files on durable storage (disk or a Docker volume). A container or VM rebuild that drops NOTIFY_DATA_DIR loses FCM registrations, delegated JWT batches, and alertSearch cursors. Back up all three files together when using WAL.
Schema (high level)
Table fcm_registrations holds one row per registered device:
- Identity:
id,user_id,device_id,fcm_token,platform - Flags:
test_mode - Timestamps:
created_at,updated_at,last_notified_at
Unique on (user_id, device_id). Indexes also exist on user_id, device_id, fcm_token, and (user_id, fcm_token).
Tables alert_authorization_batches and alert_authorization_jwts hold a user's delegated notification-JWT inventory (separate from device registration):
- Batch:
id,user_id(authenticated DID),batch_id,notify_hour_min_utc(zero-paddedHH:MMUTC, nullable),timezone(IANA name, nullable, recorded but unread),created_at - JWT:
batch_pk,sequence,day(YYYY-MM-DD),jwt,nbf,exp,status(unused/consumed),consumed_at, timestamps
Unique on (batch_pk, sequence) and on (user_id, day) for unused rows. Indexes also exist on (user_id, status), (user_id, day), and batch_pk.
Table alert_search_cursors holds one row per user DID:
endorser_after_id— last complete Endorser ULID (afterId), or nullpartner_after_at— last complete PartnerupdatedAtbound (afterDate), or nullcreated_at,updated_at
Table sms_registrations holds one row per (DID, phone) pair:
- Identity:
id,user_id,phone_e164(E.164 normalized) - Verification:
verified,code_hash(HMAC of the pending code),code_expires_at,code_attempts,last_code_sent_at,verified_at - Timestamps:
created_at,updated_at
Unique on (user_id, phone_e164). Indexes also exist on user_id, on
phone_e164 (opt-out arrives with the number, not the DID), and on
(user_id, verified).
Table sms_phone_log records every phone action: user_id, phone_e164
(nullable, nulled by DELETE), phone_hash (HMAC, never nulled), action,
result (ok / rejected / failed), detail, jwt_hash,
provider_message_id, created_at. Actions are register-requested,
code-sent, code-send-failed, verify-succeeded, verify-failed,
did-limit-blocked, did-limit-disclosed, deleted,
alert-authorization-stored, alert-authorization-deleted, alert-sent,
alert-send-failed,
recipient-not-allowed, number-blocked, number-unblocked, and opt-out.
Indexes on (user_id, created_at), (phone_hash, created_at), and
(action, created_at); the throttle counts read the second.
This is the first table in this database with directly identifying personal
data. phone_e164 is nulled on delete; a prune job for rows older than 400
days is a follow-up.
Table sms_blocked_numbers is the suppression list: phone_hash (unique, the
matching key), phone_e164, reason (opt-out / provider-opt-out / manual), detail, and
timestamps. It is the one table that deliberately retains a number after the
owner has asked to be left alone, because honoring that request means
remembering it.
Table sms_action_jwt_use holds the replay guard: jwt_hash (unique),
user_id, action, used_at. Rows older than
SMS_ACTION_JWT_MAX_AGE_SEC × 10 are pruned on each SMS scheduler pass; a token
that stale fails the freshness check anyway.
Tables sms_alert_authorization_batches / sms_alert_authorization_jwts and
sms_alert_search_cursors are column-for-column mirrors of their FCM
counterparts, holding the SMS channel's independent JWT inventory and cursor.
The schema is created automatically on startup if the database or tables do not already exist. New tables are added with CREATE TABLE IF NOT EXISTS; existing fcm_registrations rows are not migrated or altered.
Backup
Persist or back up the SQLite files under NOTIFY_DATA_DIR:
- Prefer stopping the service, then copy
notify.sqlite(and any-wal/-shmsidecars if present). - Or, while the service is running, copy all three files (
notify.sqlite,-wal,-shm) together so the backup stays consistent under WAL mode. - For Docker, mount a volume at the data directory (or set
NOTIFY_DATA_DIRto a mounted path) so registrations survive container recreation.
Production
Canonical production deployment is the Docker image: pnpm build then node dist/index.js (Dockerfile CMD). That is not the same as the development commands pnpm run dev / pnpm start, which run TypeScript through tsx and are for local development only.
Single replica
Run exactly one Node process / one production replica. AlertSearch scheduling and its in-flight overlap guard are process-local. There is no distributed scheduler lock. SQLite is a local file. Two processes will double-run AlertSearch and FCM wakeup and can corrupt or fork the database.
Firebase credentials
Working Firebase Admin credentials are required for both WAKEUP_PING and AlertSearch FCM (type: alert_search). The app initializes Firebase once at process start (src/services/firebase.ts).
Supported paths (in this order):
FIREBASE_SERVICE_ACCOUNT_JSON— inline service-account JSON (one line). The application reads this variable.- If that variable is unset or empty, Application Default Credentials. ADC may use
GOOGLE_APPLICATION_CREDENTIALS(a file path to a key JSON). The application does not readGOOGLE_APPLICATION_CREDENTIALSitself; the Google/Firebase ADC stack does.
Invalid FIREBASE_SERVICE_ACCOUNT_JSON prevents the process from starting. Missing ADC typically allows listen//health but FCM sends fail later.
Persistent SQLite
Set NOTIFY_DATA_DIR to a durable directory, or keep the Docker default /app/data on a persistent volume. The service uses:
{NOTIFY_DATA_DIR}/notify.sqlite{NOTIFY_DATA_DIR}/notify.sqlite-wal{NOTIFY_DATA_DIR}/notify.sqlite-shm
Environment checklist
| Variable / constraint | Production requirement | Default if unset |
|---|---|---|
PORT |
Optional | 3003 |
ENDORSER_URL |
Optional if using production Endorser | https://api.endorser.ch |
PARTNER_URL |
Optional if using production Partner | https://partner-api.endorser.ch |
FIREBASE_SERVICE_ACCOUNT_JSON or working ADC |
Required for FCM (wakeup and AlertSearch) | ADC if JSON unset |
NOTIFY_DATA_DIR |
Durable path (or volume on /app/data) |
./data (cwd-relative; in Docker that is /app/data) |
NODE_ENV |
Must not be test-local (that bypasses ethr JWT expiry) |
Docker image sets production |
SMS_ENABLED |
Optional; /notify-sms returns 503 SMS_DISABLED while off |
false |
SMS_CODE_SECRET |
Required when SMS_ENABLED (startup fails without it) |
None |
TWILIO_ACCOUNT_SID, TWILIO_AUTH_TOKEN |
Required to send; absent means SMS_NOT_CONFIGURED per send |
None |
TWILIO_MESSAGING_SERVICE_SID or TWILIO_FROM_NUMBER |
One of the two required to send; prefer the Messaging Service, which wins when both are set | None |
TWILIO_WEBHOOK_URL |
The public URL Twilio posts /notify-sms/inbound to; it signs that exact string |
Derived from request headers |
SMS_CODE_TTL_SEC |
Optional | 600 |
SMS_CODE_MAX_ATTEMPTS |
Optional | 5 |
SMS_ACTION_JWT_MAX_AGE_SEC |
Optional | 300 |
SMS_MAX_DIDS_PER_PHONE |
Optional | 5 |
SMS_ALERT_SEARCH_INTERVAL_MS |
Optional | 300000 |
SMS_REQUIRE_ACTION_CLAIM |
Must not be a false value in production | true |
SMS_ALLOWED_RECIPIENT_DIDS |
Leave unset in production; on a test server, set it to the DIDs that server may text | Unset (no restriction) |
SMS_BLOCKED_NUMBERS |
Optional; numbers blocked by configuration, on top of the sms_blocked_numbers table |
Empty |
SMS_DEV_ECHO_CODE |
Must be unset or false; honored only under NODE_ENV=test-local |
false |
| Replicas | One process | Not enforced in code |
| Persistent volume | Required for Docker so SQLite survives replace | None unless you pass -v |
Boolean flags (SMS_ENABLED, SMS_REQUIRE_ACTION_CLAIM, SMS_DEV_ECHO_CODE, DEBUG_ENDPOINT) accept true/1/yes/on and false/0/no/off, case-insensitive, with surrounding whitespace ignored. Any other value, including a quoted one like "true" (which docker run --env-file passes through with its quotes), leaves the flag at its default.
Rotating SMS_CODE_SECRET invalidates every pending verification code and
orphans every stored phone_hash. Rotate between deploys, not casually.
Optional: DEFAULT_ENDORSER_API_SERVER / DEFAULT_PARTNER_API_SERVER are honored only if the corresponding ENDORSER_URL / PARTNER_URL is unset (src/env.ts).
Docker (canonical)
docker build --no-cache -t notify-wakeup-api:amd-$NOTIFY_WAKEUP_API_VERSION --platform linux/amd64 .
docker run --env-file notify-wakeup-api.env -p 3003:3003 \
-v notify-wakeup-data:/app/data \
notify-wakeup-api
The image runs node dist/index.js. Mount a volume at /app/data (or set NOTIFY_DATA_DIR to another mounted path). Do not scale this container to multiple replicas.
Smoke test
Every line this service prints is prefixed with an ISO-8601 UTC timestamp
(src/util/log.ts).
/health only means the HTTP server is up. It does not prove Firebase, Endorser, Partner, SQLite durability, or AlertSearch.
- Listening:
curl -sS -o /dev/null -w "%{http_code}\n" http://127.0.0.1:3003/health(or the host/port you published). Expect200. - Health body:
curl -sS http://127.0.0.1:3003/health→{"ok":true}. - The schedulers are started from
src/index.ts(startScheduler(), thenstartAlertSearchScheduler(), thenstartSmsAlertSearchScheduler()whenSMS_ENABLED) as the process reaches* Running backend. None of them runs a pass on startup; the first pass is on the timer, and the SMS pass waits an extra 150 seconds. - AlertSearch activity in logs: look for
[AlertSearchScheduler] Pass startedand[AlertSearchScheduler] Pass completed in(orPass skipped (already in flight)). User failures log as[AlertSearchScheduler] User failed. FCM wakeup passes log[Scheduler] Pass started/[Scheduler] Pass completed in. These lines appear after the first interval, not at boot. - SMS activity in logs (when
SMS_ENABLED): look for[SmsAlertSearchScheduler] Pass started/Pass completed in. Route activity logs under[NotifySms], and rejected authorizations under[SmsActionJwt] Rejected. - SQLite location: after the process has handled a request or a scheduler pass that opens the DB, confirm
{NOTIFY_DATA_DIR}/notify.sqliteexists (Docker default:/app/data/notify.sqliteon the volume). WAL sidecars may appear while the process is running.