About SAFR database storage
SAFR uses Mongo (and open source database) to store its person and event metadata (images stored on disk). Mongo is a non-SQL database that stores data in JSON data structures.
Mongo Clustering
Mongo can be clustered so that mutliple nodes share same data and share query load. Mongo clustering technology works such that one mongo server assumes the role of primary and all other mongo nodes assume the role of secondary. The primary is responsible for all writes and also read operations. The Mongo secondaries are responsible for read operations only. All data is replicated across all nodes in a mongo cluster. Data is written to the mongo primary and replicated out to the mongo secondaries. The replication typically takes milliseconds.
Mongo Failover and elections
Should the primary instance in a mongo cluster go offline, Mongo supports a failover technology thru what is called the election process. The election process can only work if there are 2 or more secondary mongo instances. If the primary becomes unresponsive, the remaining mongo secondary instances will hold an election and one of the remaining instances will be voted as primary. That new primary instance will then take over the responsibility of writes and remaining instances will continue to function as secondaries. Should the primary instance become available again, it will assume the role of secondary.
SAFR Clustering vs Mongo Clustering
SAFR also supports clustering and has a similar notion of primary vs. secondary not following the same paradigm as Mongo primary and secondary servers. A SAFR primary server is the one responsible for obtaining and maintaining a license for the cluster and assumes the role of load balancer when performing software load balancing. Also, the primary SAFR server does not need to conincide with the mongo primary server.
By default (at installation time), the primary SAFR server also hosts the mongo primary instance. But in cases of failover, other mongo nodes can assume the role of primary and thus Mongo primary may run on a SAFR secondary. This would happen if the SAFR primary server became unavailable (either due to beiong disconnected from the network or system failure). In this case the remaining mongo instances will collaborate to elect a new primary member. Writes will then begin to occur on the newly elected primary mongo instance.
Number of nodes for failover
You can deploy any number of SAFR servers in a cluster. SAFR requires 3 nodes in order to achieve ability to have automatic failover of primary. With 2 you have redundant data but if primary goes offline, no writes will occur (only reads from the secondary). When adding more than 3, only the first three are voting members. The rest are non-voting, priority 0.
Simple vs. Redundant Joins
By default, a SAFR Server cluster will have one Mongo instance on each server in the cluster. Which nodes act as data nodes can be controlled during installation. A Simple join will add an additional server that only performs face recognition. A Redundant join will add an additional server to the cluster that performs both face recognition and data replication. Any one of the Redundantly joined SAFR servers host a Mongo primary or secondary instance.
Reads and Writes
As noted above, mongo primary is responsible for write or read operations and mongo secondaries are responsible for read only. To help reduce load of the mongo primary instance, mongo allows operations a preference to be set on read operations (secondary or primary). Primary read preference may be set in cases where the most up to date data is critical (i.e. load balancing or configuration changes). Secondary read operations are set in cases where some out-of-date data can be tolerated in order to attempt to read load on the primary (to leave it free for write operations). SAFR sets following read operations to secondary preferred:
- Covi (face matching)
- Events (search-by-face only)