Sentimus is in beta —

Sentimus

Human Sensing as a Service — Whitepaper v1

An open instrument for community-driven impact verification.

Version 1.0 · May 2026 · sentimus.co


Summary

Impact verification at the budgets where most of the world's projects actually operate is structurally broken. Third-party audits cost more than the programmes they audit. Self-reporting by the implementing organisation is the failure mode funders have learned not to trust. IoT and remote sensing answer a different question altogether — they can confirm that something physical happened, but not whether it mattered to anyone.

The people who can answer that question — the community living next to the project — already know. They have always known. They have lacked an instrument.

Sentimus is that instrument. It is the engineering of a decade-old research field — Human-as-a-Sensor, Participatory Sensing, Mobile Crowd Sensing — into a deployable platform for the specific problem of community-driven impact verification. The methodology is published. The scoring formulas are mathematically conservative by design. The architecture is open, provider-agnostic, and offline-first by necessity.

This paper situates Sentimus inside the academic literature it inherits, sets out the standards vacuum it operates in, derives the math that lets fifteen trustworthy responses carry the weight of a hundred unverified ones, and explains how a smaller prior project — Sentlog — provides the cryptographic substrate that makes any of this defensible.


1. The standardisation vacuum

Sentimus exists because of a gap that the standards bodies have not yet closed.

For machine sensors, the standards are mature:

  • ISO 19156 (Observations & Measurements)
  • OGC SensorThings API
  • W3C Web of Things
  • SenML (RFC 8428) and IPSO Smart Objects

Each of these defines an observation model — calibration profile, error characteristics, measurement range, output schema — for a device producing a reading. None of them has any vocabulary for a human producing one. There is no equivalent of ISO 19156 for human observers. As the trust research notes, "no ISO standard addresses the 'human as sensor' paradigm."

For social impact, the standards define what to measure, not how to trust the collector:

  • IRIS+ (GIIN) — standardised metrics aligned with the SDGs.
  • Impact Genome Registry — 132+ standardised social outcomes; binary expert verification, no community voice.
  • Impact Evaluation Standard — structured evaluation protocols.

For digital MRV (measurement, reporting, verification), credible options exist for environmental impact — Pachama, Sylvera, Perennial — all built on satellite imagery and remote sensing. Every one of them is environmental. None of them verify social impact, because social impact is not visible from a satellite.

In between sits a $50k–$150k Gold Standard or Verra audit, designed for programmes whose total budgets exceed seven figures. Most development projects budget between $1,000 and $250,000. Mathematically, the audit floor structurally excludes the very category of project for which independent verification matters most.

The seven gaps the literature identifies — no human sensor model, no confidence-and-context metadata, no cross-domain interoperability, no standard for legitimate disagreement, no calibration protocol, no reconciliation of standardisation with community sovereignty, no quality-reputation-incentive integration — are not seven separate problems. They are the same problem in seven coordinates: there is no engineered instrument for trustworthy community sensing.

Sentimus is one.


2. The academic lineage Sentimus inherits

This is not a new field. We are engineering inside a research tradition with a decade of foundational work and several decades of methodological scaffolding behind it.

Human-as-a-Sensor (HaaS). Dong Wang, Tarek Abdelzaher and Lance Kaplan's Social Sensing: Building Reliable Systems on Unreliable Data (Elsevier, 2015) established the substrate. Humans are sensing nodes in an information network. The textbook problem — extracting reliable information from data collected from largely unknown and possibly unreliable sources — is the problem Sentimus solves at the verification layer.

Participatory Sensing. Coined at UCLA's Center for Embedded Networked Sensing (CENS) in the mid-2000s by Deborah Estrin, Jeffrey Burke and collaborators. The foundational 2006 Woodrow Wilson Center paper formalised campaign-based organisation, multi-scale participation (personal, community, urban), data quality through redundancy, and privacy by design.

Mobile Crowd Sensing and Computing (MCSC). Bin Guo's lineage paper traces the field from CENS to MCSC. The ACM Computing Surveys (2015) survey distinguishes explicit from implicit participation and personal from public sensing.

Citizen Science. Zooniverse manages quality through redundancy and statistical aggregation. The CrowdTruth framework (VU Amsterdam) makes the case that disagreement is not noise to be eliminated but informative signal to be modelled.

The Delphi Method (Helmer & Dalkey, RAND, 1950s). A structured, iterative, anonymous, statistically aggregated process for eliciting expert judgement. The four design principles — anonymity, iteration, controlled feedback, statistical aggregation — appear, transformed, in the Sentimus pulse model.

Sensemaking. Karl Weick's Sensemaking in Organizations (Sage, 1995) names seven properties of how groups construct meaning — retrospective, social, ongoing, driven by plausibility rather than accuracy. Brenda Dervin's situation-gap-uses methodology reframes information-seeking as gap-bridging. Both inform how Sentimus presents distributions rather than single scores.

Participatory Monitoring & Evaluation. Robert Chambers' Whose Reality Counts? (IDS Sussex, 1997) identifies six biases that distort outsiders' understanding: spatial, project, person, seasonal, diplomatic, professional. Davies & Dart's Most Significant Change technique (2005) and David Bonbright's Constituent Voice methodology (Keystone Accountability) ground the practice. Sentimus operationalises Chambers' bias inventory through proximity weighting and dual-perspective scoring.

Wisdom of Crowds. Surowiecki (2004) gives the four conditions for collective intelligence; Condorcet's Jury Theorem (1785) gives the math. Both appear below.

We name this lineage explicitly for two reasons. First, because building inside a research tradition is honest engineering, not novelty theatre. Second, because every counter-argument to community sensing has been raised, examined, and partially answered somewhere in this literature. We do not need to rediscover the problems. We need to engineer around them.


3. Why a human is the right sensor

A physical sensor measures one variable precisely until it breaks. A human sensor measures meaning, context, and change — wherever people are.

This is not metaphorical. The properties of the two instruments are structurally different:

Physical sensor Human sensor
Captures meaning Never Always
Covers social impact Impossible Native
Adaptive Measures only what it was designed for Notices what nobody predicted
Cost at scale Linear ($ per device) Logarithmic (network effect)
Works offline Needs backhaul Needs nothing (syncs later)
Longitudinal Until hardware fails As long as people live there
Self-replicating No Yes — people talk to people
Survives infrastructure failure No Yes

An IoT water-flow meter on a borehole reports 12 L/min flowed yesterday. It cannot report whether the walk for water got shorter, whether children stopped getting sick, whether the maintenance committee actually meets, whether the queue is fair, whether anyone's life is different in any way that matters. Those questions are not in the borehole's measurement range. They are in the community's.

Human sensing is not the fallback for places where IoT cannot reach. It is the primary instrument for the questions IoT structurally cannot ask. Where physical sensors fit, they should be used; where both exist, they corroborate. Resilience to infrastructure failure is a secondary consequence of the human-sensor model — useful, but not the headline.


4. The trust problem — and the math that bounds it

The intuitive objection to community testimony is fair: individuals are biased, mistaken, sometimes coached, occasionally gaming the system. Every one of these criticisms is true at small n.

At scale, under specific conditions, they cancel.

Surowiecki's Wisdom of Crowds (2004) identifies the four conditions:

  1. Diversity of opinion — each person holds private information
  2. Independence — opinions are not influenced by others'
  3. Decentralisation — people draw on local knowledge
  4. Aggregation — a mechanism turns private judgments into a collective signal

When the conditions are met, Condorcet's Jury Theorem (1785) gives the math: if each respondent has a better-than-chance probability of being correct, the probability that the majority is correct approaches 1 as n increases.

Respondents (n) Individual accuracy Majority correct
10 60% ~63%
30 60% ~85%
100 60% ~97%
300 60% ~99.9%

The implication is uncomfortable for skeptics of community testimony: at n = 10, a single gaming respondent is 10% of the signal; at n = 100, that respondent is noise.

But the engineering target is not n = 100. The Sentimus model deliberately reduces the required sample size through three reinforcing mechanisms drawn directly from the modern crowd-sensing literature:

  1. Machine evidence — on-device sensors corroborate human testimony.
  2. Proximity weighting — a beneficiary's evidence counts more than an outsider's hearsay.
  3. Respondent reliability over time — known-reliable respondents outweigh unknowns.

With all three in place, n ≈ 15 with corroboration produces trust comparable to n ≈ 100 without it. The aggregation math that delivers this is explicit, published, and below.


5. The four-dimensional evidence model

Every Sentimus response captures four orthogonal signals. None is sufficient alone; all four together resist the failure modes any one of them would exhibit on its own.

Hard evidenceDid it happen? A direct yes / no / don't-know on the claim. Drives the confidence score.

Soft evidenceDid it matter? A visual scale (face-anchored continuous slider, validated by WHO for low-literacy cross-cultural settings) or a four-option before/after comparison anchored to the respondent's own baseline. Drives the sentiment score. This is the dimension physical sensors cannot reach at all.

Machine evidenceCan the phone corroborate? On-device ML photo classification, EXIF metadata, GPS coordinates, GPS movement trail, accelerometer motion signature, barometer altitude, ambient light, timestamp. Drives the authenticity score.

ProximityHow close are you to the impact? Beneficiary → household → neighbour → community → outsider. Drives both a per-response weight multiplier and a perspective separation — inner circle (lived experience, depth) versus outer circle (external observation, breadth).

5.1 The authenticity score

Machine evidence produces an authenticity score per response — a measure of "was this person physically present, did they photograph something relevant, and do the sensor readings match reality?":

authenticity = w_photo · photo_relevance
             + w_gps   · gps_match
             + w_time  · time_match
             + w_env   · environment_consistency

Default weights: w_photo = 0.40, w_gps = 0.30, w_time = 0.15, w_env = 0.15.

photo_relevance is the ML classifier's confidence that the photograph contains the activity type claimed (water infrastructure, cookstove, training group, etc.) using per-activity TFLite models that run in under 500 ms on a 2020-era Android. gps_match is 1.0 within the claim radius, decaying with distance. time_match is 1.0 within the pulse window. environment_consistency is a composite of indoor/outdoor light level, motion signature, and altitude against known elevation. No image data leaves the device for inference; classification happens at capture time.

5.2 The proximity scale and dual perspective

The proximity weights are not arbitrary. They are calibrated to the depth versus breadth of the observation each tier can produce:

beneficiary  1.00   "I use the borehole daily"
household    0.85   "My wife fetches water there"
neighbour    0.60   "I see people using it"
community    0.35   "I heard about it"
outsider     0.15   "I don't live here"

A beneficiary provides depth — actual change in daily life. A neighbour provides breadth — relative comparison and visibility. Both are valid; both belong in the evidence pool; their weights differ because the kinds of evidence they can produce differ.

These five tiers split into two perspectives:

  • Inner circle (beneficiary + household) — lived experience; depth of impact
  • Outer circle (neighbour + community + outsider) — external perception; breadth of visibility

Sentimus does not collapse the two into a single number. It reports them separately. The pattern of agreement or disagreement between inner and outer circles is itself informative:

Inner Outer Interpretation
High High Impact is both felt and visible — strong signal
High Low Real but invisible impact (e.g. health training — beneficiaries feel better, neighbours can't see it)
Low High Surface-level impact (built but not accessible to beneficiaries)
Low Low Weak claim

This directly answers the trust literature's open problem of "legitimate disagreement as data, not noise" (the CrowdTruth thesis). For genuinely subjective phenomena, the model does not impose a single ground truth; it reports the distribution and lets the consumer interpret divergence.

5.3 The combined scoring model

Per response:
  hard_score    ∈ {1, 0, null}
  soft_score    ∈ [0, 1]              visual_scale or before/after
  authenticity  ∈ [0.1, 1]            (floor at 0.1 for no-evidence responses)
  weight        = proximity × authenticity

Per perspective (inner, outer):
  confidence = Σ(hard_i · weight_i) / Σ(weight_i)
  sentiment  = Σ(soft_i · weight_i) / Σ(weight_i)

Per claim:
  confidence  = weighted_mean(inner_confidence, outer_confidence)
  sentiment   = weighted_mean(inner_sentiment, outer_sentiment)
  convergence = 1 − |inner_combined − outer_combined|
  authenticity_avg = mean(per-response authenticity)

  Final confidence capped at 0.85.

Authenticity acts as a multiplier on weight, not as a separately published score. A response with low authenticity (no photo, GPS mismatch) counts less; a response with high authenticity counts more. Machine evidence naturally amplifies trusted responses and dampens suspicious ones inside the same scoring pipeline.

5.4 Effective sample size

The aggregation produces an effective n that captures how much trustworthy signal exists, distinct from raw response count:

effective_n = Σ (authenticity_i × reliability_i × proximity_weight_i)

A pulse with 15 responses where average authenticity is 0.85, average reliability is 0.7, and average proximity weight is 0.75 produces an effective n of ~6.7 — but those are 6.7 fully trusted data points, equivalent to roughly 30 unverified responses. The model trades raw count for verified count, honestly.

5.5 Conservatism is in the math

Two parameters are non-negotiable, and they are not tunable per claim:

  • Confidence cap: 0.85. Community surveys are inherently imperfect; we refuse to publish certainty we cannot mechanically justify.
  • Raw signal discount: 50%. Applied before any of the above. Sentimus would rather under-report a real impact than over-report one.

Both are tested. Both are visible in the source. Both apply regardless of the funder, the org, or the claim. Conservatism is not a tone choice — it is structural in the formulas.


6. Anti-gaming — the cost-to-fake chain

The hardest problem in any self-reported system is fabrication. The literature is clear that a lying human produces data that looks plausible, while a broken thermometer produces data that looks obviously wrong (Wang et al., 2015 §3).

The four-dimensional model raises the cost of a successful fabrication to the point where, for most claim types, faking the response is structurally more expensive than producing the underlying real evidence.

To successfully game a single Sentimus response, a respondent must simultaneously:

  1. Answer the hard-evidence questions plausibly — easy
  2. Provide a consistent sentiment score — easy
  3. Select a credible proximity level — easy
  4. Be physically present at the GPS coordinates of the claimrequires travel
  5. Photograph something the ML classifier accepts as the correct activity typerequires the thing to exist
  6. Produce environmental sensor readings consistent with field presence (accelerometer signature, barometer altitude, ambient light) — hard to synthesise
  7. Do all of the above within the pulse windowtime pressure

Any one failure flags the response as low-authenticity and cuts its weight inside the scoring pipeline. Coordinated gaming (multiple fabricated respondents) compounds the requirement: multiple devices at the correct location, multiple photographs from different angles, diverse proximity claims consistent with the GPS data, response timing that does not look mechanically uniform, sensor data that does not look copied.

The residual vulnerability is candid: a respondent physically present at the site with the right equipment can game the system. But if they are physically at the borehole, photographing the borehole, with motion and altitude readings consistent with the borehole — the borehole exists. The gaming has become indistinguishable from genuine verification. The effort to fake has exceeded the effort to be honest.

This is what the trust research means by raising the cost of fabrication. We do not eliminate the possibility. We engineer the economics.


7. The seven critical gaps — and what the model resolves

The trust literature identifies seven structural gaps that no existing platform closes. Sentimus does not solve all seven; it makes measurable progress on each.

Gap Status
1. No human sensor model Addressed. Proximity + reliability + machine evidence constitutes a de facto model. Not ISO-standardised; implementable today.
2. No confidence/context metadata Resolved. EvidenceBreakdown encodes confidence, sentiment, proximity, authenticity, convergence and perspective split per observation.
3. Cross-domain interoperability Partial. @sentimus/shared is domain-agnostic; the model works for any activity type. No formal interoperability with health/environment/urban domains yet.
4. No standard for legitimate disagreement Resolved. Dual-perspective preserves disagreement as informative signal. Convergence quantifies agreement without forcing consensus. This is the model's strongest theoretical contribution.
5. No calibration protocol Partial. Proximity-as-calibration plus visual/pictorial response formats reduce calibration needs. Per-respondent reliability tracking deferred to a later phase.
6. Standardisation vs. sovereignty Addressed by design. The model defines how to structure evidence, not what counts as impact. Communities define their own questions and success criteria inside a common output format.
7. Quality-reputation-incentive integration Partial. Quality → authenticity → weight in confidence. Reputation → reliability tracking (deferred). Reward → quality-proportional distribution. Full integration deferred.

What the model does not solve is named below, in §10. We hold ourselves to surfacing the residuals as visibly as the resolutions.


8. The trust mechanisms drawn from the modern literature

Three mechanisms from the contemporary crowd-sensing literature are integrated directly into the scoring model:

Dawid-Skene (1979, with 2024 Human-AI extension). Per-respondent confusion matrices estimated via Expectation-Maximisation, alternating between truth estimates and reliability estimates. Expert annotators automatically receive more weight; spammers are downweighted. The 2024 extension permits AI classifier probability distributions to enter the same aggregation. In Sentimus, this manifests as the reliability_i term in effective_n and is used to update per-respondent priors across pulses.

Bayesian Truth Serum (Prelec, Science 2004; validated at scale, 2024–25). Adds an optional meta-question: "What do you think most people in your community will answer?" The information score:

InformationScore(k) = log( actual_frequency_k / predicted_frequency_k )

rewards "surprisingly common" answers — responses more frequent than the respondent themself predicted. Truthful reporting is a Bayesian Nash equilibrium under BTS. Empirically, BTS makes n = 30 behave like n = 100 by extracting more signal per response.

CROWDLAB (Goh, Tkachenko, Mueller — NeurIPS Workshop 2022). Treats a trained classifier as an additional "annotator" alongside humans, outperforming Dawid-Skene, GLAD and majority vote on real datasets. In Sentimus this is the structural basis for treating the edge-AI photo classifier and the human respondent as parallel evidence channels in the consensus computation, rather than treating ML as a separate verification stage.

Reading the trust research and the methodology in parallel: every formula above maps to a published algorithm with measured empirical performance. None of this is invented for the whitepaper.


9. Lineage — from Sentlog to Sentimus

Sentimus did not start as an impact-verification platform. It started as a question, asked by an earlier and smaller project called Sentlog: can a phone produce evidence trustworthy enough to be cited?

Sentlog is a device-first evidence SDK. Its scope is narrow and deep. It captures a short voice interaction — optionally enriched with motion and facial-affect signals — and produces a multi-dimensional presence assessment, emitted as a cryptographically signed W3C Verifiable Credentials 2.0 envelope. Capture and inference happen entirely on-device. The default posture is inert: nothing transmits without explicit integrator configuration.

The technical commitments Sentimus inherits from Sentlog are specific:

  • W3C VC 2.0 envelope with Ed25519 or ECDSA-P256 signatures
  • Model version pinned into the credential (so the inference run is auditable years later)
  • Perceptual audio fingerprint carried as a tamper-evident anchor
  • Optional RFC 3161 timestamping for legal-grade temporal proof
  • Universal deployment — native iOS, native Android, Capacitor PWAs, React Native — on a shared C++ core with platform bridges
  • Composed signals exposed separately — liveness, engagement, direction, presence proper, optional facial affect. The integrator defines the verdict logic; the SDK refuses to.

Sentlog answers the substrate question. Yes — a phone, on its own, can produce evidence that is signed, attributable, replayable, refusable, and tamper-evident. This answer is necessary for any honest community-sensing platform to exist. It is not, on its own, sufficient.

Evidence on one phone is not impact measurement. A single signed response is a data point; a community signal is a different mathematical object entirely, governed by aggregation rules, proximity weights, reliability priors, and a published methodology. Sentimus is the network and aggregation layer that operates on top of the cryptographic substrate Sentlog establishes.

Both projects share four operating principles — device-first, open, signed, composable — and the same refusal to centralise what belongs to the people producing the data. Sentlog proves the unit. Sentimus is the network.


10. What remains genuinely hard

Three problems survive the four-dimensional model. We name them because intellectual honesty requires it, and because the credibility of every claim above depends on not over-claiming here.

10.1 Coordinated organisational fraud

An organisation that controls the entire survey process — writes the questions, distributes the response tokens, coaches the respondents, and submits the responses — can game the system even with machine evidence. The structural problem is that the claim-maker and the evidence-collector are the same entity. There is no independent observer.

Available mitigations: peer challenges from other organisations operating in the same region; statistical anomaly detection (suspiciously uniform responses, identical response times, clustered device fingerprints); third-party replay — a funder can independently trigger a new pulse and distribute tokens through their own channels.

Unavailable without external input: an independent verification oracle. This is the structural limit of any self-reported system, not unique to Sentimus.

10.2 Sparsity in low-population contexts

A project serving eight households in a remote area. There are not thirty respondents to survey. There may be eight. Condorcet's theorem requires scale; machine evidence helps but does not fully compensate below a minimum threshold.

Available mitigations: effective-n amplification through machine evidence (eight responses with high authenticity ≈ twelve to fifteen effective); proportionally lower confidence cap, reporting the appropriate uncertainty rather than fabricating it; longitudinal evidence — the same eight people surveyed three times over twelve months produces twenty-four data points.

The honest answer is that below MIN_TOTAL_RESPONSES = 5 the platform reports insufficient_data and refuses to publish a score. We cannot manufacture evidence that does not exist.

10.3 The cold start

The first pulse for a new organisation in a new region has no respondent history, no reliability data, no anomaly-detection baseline. Every safeguard that depends on accumulated history is absent on day one.

Available mitigations: machine evidence and proximity weighting work from the first response. Minimum-n with high authenticity is required for first pulses. Confidence is published with a "maturity" flag indicating how much historical context exists.

The model is weakest on its first use and strongest after months of longitudinal data. This is inherent to any learning system and we surface it explicitly rather than smoothing it away.


11. How Sentimus operates

The end-to-end flow is deliberately short.

1. Claim and pulse. The implementing organisation publishes what is being tested, how, and against which funding decision. The claim is versioned, public, and tied to the programme it belongs to.

2. Community responds. Phones in the field — anonymous, offline-capable, with photo, GPS and on-device classification attached where the respondent permits. The PWA installs on whatever device the respondent already carries; no app-store gatekeeping, no new hardware, no enterprise platform to learn.

3. Evidence combines. Hard, soft, machine and proximity signals are aggregated under the published scoring rules. The 50% discount and 0.85 confidence cap apply in every case.

4. Signal exports. Scores flow out through a REST API, webhooks, and an MCP server into whatever reporting tools the organisation already uses. There is no platform dashboard the organisation must log into and no exit cost if it later decides to leave. The data belongs to the organisation that produced it.

Several characteristics of the platform are non-negotiable:

  • Offline-first. The PWA works without connection; responses queue locally and sync when the network returns.
  • Anonymous by design. Respondents access surveys via single-use tokens — no account, no profile, no permanent identifier any party can use to retaliate.
  • Conservative by mechanism. The discount and the confidence cap are not tunable.
  • Inspectable. Methodology, scoring formulas, brand tokens, API surface, and platform architecture are all written down. Nothing critical is a trade secret.

12. Operating principles

Sentimus is built around five principles we hold prior to features:

Open Architect. Built on an open-source stack. Public methodology. Open architecture. No vendor lock-in, no walled gardens, no proprietary protocols. Where competitors guard their methods as moats, we publish ours.

Provider-agnostic. API keys per organisation, webhooks for integration, MCP for AI tooling, adapters for the systems organisations already operate. Sentimus is the substrate; the dashboards and reports live wherever the organisation already keeps its work.

Infrastructure-only. Sentimus never custodies funds, never holds beneficiary identity, and never positions itself between the organisation and its community. We provide the instrument. The organisation operates the process.

Tools, not managed service. Sentimus does not run pilots on the organisation's behalf. The team that knows the community is the team that owns the engagement. We provide the instrument and get out of the way.

Conservative by mechanism. Wherever a choice exists between flattering the data and bounding it honestly, we bound it. The credibility of the entire platform rests on the willingness to publish a low score when the evidence supports a low score.


13. Where Sentimus fits — and where it does not

Sentimus is built for projects where reporting needs community input and external auditing is impractical: small and mid-size programmes, tight budgets, limited timelines, and the situations where a six-figure third-party evaluation is not on the table and the claim still has to be honestly assessed.

It works as the primary verification layer for those projects. It works as an auxiliary layer alongside existing M&E flows. It does not replace physical sensors where they fit, and it does not replace trained third-party auditors for high-stakes regulated reporting. Where those instruments belong, they belong; Sentimus adds the community-side dimension nobody else is measuring at all.

Reporting that requires a regulator-recognised auditor — carbon credit certification at scale, financial audit, clinical trial endpoints — is outside our scope by design. Sentimus is the right tool for the very large class of programmes that have no such apparatus available and no realistic path to procuring one.


14. The instrument is open. The work is not solitary.

Sentimus is in closed beta. A small group of pilot projects is starting to run real community assessments and shape the platform with us. The methodology, this whitepaper, and the architectural documentation are public. The repository will follow.

If your organisation runs a project where the community side of impact is currently invisible — where you know the work is changing things and have no credible way to publish that — we would like to hear from you.

  • Apply to the closed beta: sentimus.co
  • Read the evidence methodology: docs/methodology.md
  • Read the trust research synthesis: docs/trust.md
  • Read the trust counter-analysis: docs/trust-counter-analysis.md
  • Write to us: [email protected]

Selected references

Foundational. Wang, Abdelzaher & Kaplan, Social Sensing (Elsevier, 2015) · Burke, Estrin, Hansen, Goldman et al., "Participatory Sensing" (Woodrow Wilson Center, 2006) · Surowiecki, The Wisdom of Crowds (Doubleday, 2004) · Weick, Sensemaking in Organizations (Sage, 1995) · Chambers, Whose Reality Counts? (Practical Action, 1997) · Davies & Dart, "The Most Significant Change Technique" (2005).

Trust mechanisms. Prelec, "A Bayesian Truth Serum" (Science, 2004) · Prelec, Seung & McCoy, "Surprisingly Popular" (Nature, 2017) · Dawid & Skene, "Maximum Likelihood Estimation of Observer Error-Rates" (1979) · Goh, Tkachenko & Mueller, "CROWDLAB" (NeurIPS Workshop, 2022) · Whitehill et al., "GLAD" (NeurIPS, 2009) · Miller, Resnick & Zeckhauser, "Peer Prediction" (2005) · Brier, "Verification of Forecasts" (1950).

Standards landscape. ISO 19156 (Observations & Measurements) · OGC SensorThings API · W3C Web of Things · SenML (RFC 8428) · IRIS+ (GIIN) · Impact Genome Registry · ACL 2021, "Empirical Approach to Inter-Rater Reliability."

A complete bibliography lives in docs/trust.md and the methodology in docs/methodology.md.


Sentimus — from Latin sentire, to sense; first-person plural, "we sense."

The instrument is the community. The platform just carries the signal.