redict

mirror of https://codeberg.org/redict/redict.git synced 2025-01-24 00:59:02 -05:00

Author	SHA1	Message	Date
antirez	9dfe426fc8	Sentinel: HELLO processing refactored into sentinelProcessHelloMessage().	2014-03-14 11:07:42 +01:00
Jan-Erik Rediger	5f5118bdad	Small typo fixed	2014-03-05 00:41:02 +01:00
antirez	47750998a6	Sentinel: more aggressive failover start desynchronization. Sentinel needs to avoid split brain conditions due to multiple sentinels trying to get voted at the exact same time. So far some desynchronization was provided by fluctuating server.hz, that is the frequency of the timer function call. However the desynchonization provided in this way was not enough when using many Sentinel instances, especially when a large quorum value is used in order to force a greater degree of agreement (more than N/2+1). It was verified that it was likely to trigger a split brain condition, forcing the system to try again after a timeout. Usually the system will succeed after a few retries, but this is not optimal. This commit desynchronizes instances in a more effective way to make it likely that the first attempt will be successful.	2014-03-04 17:09:36 +01:00
antirez	b15411df98	Sentinel: log quorum with +monitor event.	2014-02-24 17:10:20 +01:00
antirez	6b373edb77	Sentinel: generate +monitor events at startup.	2014-02-24 16:33:55 +01:00
antirez	3b7a757468	Sentinel: log +monitor and +set events. Now that we have a runtime configuration system, it is very important to be able to log how the Sentinel configuration changes over time because of API calls.	2014-02-24 16:33:43 +01:00
antirez	25cebf7285	Sentinel: added missing exit(1) after checking for config file.	2014-02-24 16:22:52 +01:00
antirez	b1c1386374	Sentinel: IDONTKNOW error removed. This error was conceived for the older version of Sentinel that worked via master redirection and that was not able to get configuration updates from other Sentinels via the Pub/Sub channel of masters or slaves. This reply does not make sense today, every Sentinel should reply with the best information it has currently. The error will make even more sense in the future since the plan is to allow Sentinels to update the configuration of other Sentinels via gossip with a direct chat without the prerequisite that they have at least a monitored instance in common.	2014-02-22 17:34:46 +01:00
antirez	7d7b3810e7	Sentinel: report instances role switch events. This is useful mostly for debugging of issues.	2014-02-20 12:13:52 +01:00
antirez	7cec9e48ce	Sentinel: SENTINEL_SLAVE_RECONF_RETRY_PERIOD -> RECONF_TIMEOUT Rename define to match the new meaning.	2014-02-18 10:27:38 +01:00
antirez	18b8bad53c	Sentinel: fix slave promotion timeout. If we can't reconfigure a slave in time during failover, go forward as anyway the slave will be fixed by Sentinels in the future, once they detect it is misconfigured. Otherwise a failover in progress may never terminate if for some reason the slave is uncapable to sync with the master while at the same time it is not disconnected.	2014-02-18 08:50:57 +01:00
antirez	e1b77b61f3	Sentinel: better specify startup errors due to config file. Now it logs the file name if it is not accessible. Also there is a different error for the missing config file case, and for the non writable file case.	2014-02-17 16:44:49 +01:00
antirez	2d6eb68993	Sentinel: allow SHUTDOWN command in Sentinel mode.	2014-02-07 11:22:24 +01:00
antirez	3ff1bb4b2e	Sentinel: check arity for SENTINEL MASTER command. This fixes issue #1530.	2014-01-31 10:13:38 +01:00
antirez	d5763dceaf	SENTINEL SET master quorum implemented.	2014-01-14 09:23:26 +01:00
antirez	fe86f890b0	SENTINEL SET: error on bad option name + flush config on error.	2014-01-13 11:55:57 +01:00
antirez	f822516e43	SENTINEL SET implemented. The new command allows to change master-specific configurations at runtime. All the settable parameters can be retrivied via the SENTINEL MASTER command, so there is no equivalent "GET" command.	2014-01-13 11:53:29 +01:00
antirez	3cdcaff069	Sentinel: fix wrong arity error message.	2014-01-13 11:05:13 +01:00
antirez	964f6b17e9	Sentinel: SENTINEL REMOVE command added. The command totally removes a monitored master.	2014-01-10 15:39:36 +01:00
antirez	cf2835519e	Sentinel: releaseSentinelRedisInstance() top comment fixed. The claim about unlinking the instance from the connected hash tables was the opposite of the reality. Also the current actual behavior is safer in most cases, so it is better to manually unlink when needed.	2014-01-10 15:33:42 +01:00
antirez	9d0f46c6f5	Sentinel: flush config on disk when new master is added.	2014-01-10 15:22:06 +01:00
antirez	39f9f449b0	Sentinel: SENTINEL MONITOR command implemented. It allows to add new masters to monitor at runtime.	2014-01-10 15:18:24 +01:00
antirez	c42e4bd0b6	Sentinel: added SENTINEL MASTER <name> command. With SENTINEL MASTERS it was already possible to list all the configured masters, but not a specific one.	2014-01-10 14:41:52 +01:00
antirez	2bb9cd464e	Add all the configurable fields to addReplySentinelRedisInstance(). Note: the auth password with the master is voluntarily not exposed.	2014-01-10 14:31:41 +01:00
antirez	5a7d04ee7b	Trip comment to 80 cols in SentinelCommand().	2014-01-10 14:13:04 +01:00
antirez	5320148883	Sentinel: dead code removed.	2013-12-13 11:01:13 +01:00
antirez	2eb781b35b	dict.c: added optional callback to dictEmpty(). Redis hash table implementation has many non-blocking features like incremental rehashing, however while deleting a large hash table there was no way to have a callback called to do some incremental work. This commit adds this support, as an optiona callback argument to dictEmpty() that is currently called at a fixed interval (one time every 65k deletions).	2013-12-10 18:46:24 +01:00
antirez	c590549e40	Sentinel: fix reported role info sampling. The way the role change was recoded was not sane and too much convoluted, causing the role information to be not always updated. This commit fixes issue #1445.	2013-12-06 12:46:56 +01:00
antirez	2b414a4b5f	Sentinel: fix reported role fields when master is reset. When there is a master address switch, the reported role must be set to master so that we have a chance to re-sample the INFO output to check if the new address is reporting the right role. Otherwise if the role was wrong, it will be sensed as wrong even after the address switch, and for enough time according to the role change time, for Sentinel consider the master SDOWN. This fixes isue #1446, that describes the effects of this bug in practice.	2013-12-06 11:37:46 +01:00
antirez	11e81a1e9a	Fixed grammar: before H the article is a, not an.	2013-12-05 16:35:32 +01:00
antirez	f80cf7363a	Sentinel: don't write HZ when flushing config. See issue #1419.	2013-12-02 15:56:10 +01:00
antirez	dffebbc904	Sentinel: better time desynchronization. Sentinels are now desynchronized in a better way changing the time handler frequency between 10 and 20 HZ. This way on average a desynchronization of 25 milliesconds is produced that should be larger enough compared to network latency, avoiding most split-brain condition during the vote. Now that the clocks are desynchronized, to have larger random delays when performing operations can be easily achieved in the following way. Take as example the function that starts the failover, that is called with a frequency between 10 and 20 HZ and will start the failover every time there are the conditions. By just adding as an additional condition something like rand()%4 == 0, we can amplify the desynchronization between Sentinel instances easily. See issue #1419.	2013-12-02 12:29:42 +01:00
antirez	0addf8aff1	Sentinel: log vote received from other Sentinels.	2013-11-28 15:23:46 +01:00
huangz1990	86a540a66e	fix a bug in sentinel.c about pub/sub link	2013-11-26 19:55:51 +08:00
antirez	6f4fd55762	Sentinel: fixes inverted strcmp() test preventing config updates. The result of this one-char bug was pretty serious, if the new master had the same port of the previous master, but just a different IP address, non-leader Sentinels would not be able to recognize the configuration change. This commit fixes issue #1394. Many thanks to @shanemadden that reported the bug and helped investigating it.	2013-11-25 10:59:53 +01:00
antirez	8d547ebd56	Sentinel: fix type specifier for Hello msg generation. This fixes issue #1395.	2013-11-25 10:24:34 +01:00
antirez	cc6053681f	Sentinel: different comments updated to new implementation.	2013-11-21 16:22:59 +01:00
antirez	685e79998c	Sentinel: cleanup around SENTINEL_INFO_VALIDITY_TIME.	2013-11-21 16:05:41 +01:00
antirez	489d889726	Sentinel: removed mem leak and useless code.	2013-11-21 15:43:55 +01:00
antirez	f55ad3038f	Sentinel: manual failover works again.	2013-11-21 12:39:47 +01:00
antirez	297de1ab26	Sentinel: test for writable config file. This commit introduces a funciton called when Sentinel is ready for normal operations to avoid putting Sentinel specific stuff in redis.c.	2013-11-21 12:28:15 +01:00
antirez	d920177f8d	Sentinel: check for disconnected links in sentinelSendHello(). Does not fix any bug as the test is performed by the caller, but better to have the check.	2013-11-21 11:35:50 +01:00
antirez	8810167d13	Sentinel: Hello message sending code refactored.	2013-11-21 11:31:06 +01:00
antirez	0101c2bcfe	Sentinel: select slave with best (greater) replication offset.	2013-11-20 16:05:36 +01:00
antirez	a6ebd910d8	Sentinel: take the replication offset in slaves state.	2013-11-20 15:53:21 +01:00
antirez	37a51a2568	Sentinel: distinguish between is-master-down-by-addr requests. Some are just to know if the master is down, and in this case the runid in the request is set to "*", others are actually in order to seek for a vote and get elected. In the latter case the runid is set to the runid of the instance seeking for the vote.	2013-11-19 16:50:04 +01:00
antirez	b22d1beea0	Sentinel: various fixes to leader election implementation.	2013-11-19 16:20:42 +01:00
antirez	1f9728cb20	Sentinel: failover script execution fixed.	2013-11-19 12:34:46 +01:00
antirez	90635488ce	Sentinel: no longer used defines removed.	2013-11-19 11:24:36 +01:00
antirez	0a35f65301	Sentinel: when writing config on disk, remember sentinels runid.	2013-11-19 11:11:43 +01:00
antirez	5450833d02	Sentinel: arity of known-sentinel/slave is 4 not 3.	2013-11-19 11:03:47 +01:00
antirez	b8a94463b7	Sentinel: rewriteConfigSentinelOption() sub-iterators var typo fixed.	2013-11-19 10:59:50 +01:00
antirez	16237d78c8	Sentinel: call sentinelFlushConfig() to persist state when needed. Also the sentinel configuration rewriting was modified in order to account for failover in progress, where we need to provide the promoted slave address as master address, and the old master address as one of the slaves address.	2013-11-19 10:55:43 +01:00
antirez	e257ab2bfe	Sentinel: sentinelFlushConfig() to CONFIG REWRITE + fsync.	2013-11-19 10:13:04 +01:00
antirez	5998769c28	Sentinel: CONFIG REWRITE support for Sentinel config.	2013-11-19 09:48:12 +01:00
antirez	47df12d5d9	Sentinel: can-failover option removed, many comments fixed.	2013-11-19 09:28:47 +01:00
antirez	232cdb95ab	Sentinel: added config options useful to take state on config rewrite. We'll use CONFIG REWRITE (internally) in order to store the new configuration of a Sentinel after the internal state changes. In order to do so, we need configuration options (that usually the user will not touch at all) about config epoch of the master, Sentinels and Slaves known for this master, and so forth.	2013-11-18 16:03:03 +01:00
antirez	3a374b0511	Sentinel: failover abort function simplified.	2013-11-18 11:43:35 +01:00
antirez	e0750acf11	Sentinel: slaves reconfig delay modified. The time Sentinel waits since the slave is detected to be configured to the wrong master, before reconfiguring it, is now the failover_timeout time as this makes more sense in order to give the Sentinel performing the failover enoung time to reconfigure the slaves slowly (if required by the configuration). Also we now PUBLISH more frequently the new configuraiton as this allows to switch the reapprearing master back to slave faster.	2013-11-18 11:37:24 +01:00
antirez	83316f515c	Sentinel: failover restart time is now multiple of failover timeout. Also defaulf failover timeout changed to 3 minutes as the failover is a fairly fast procedure most of the times, unless there are a very big number of slaves and the user picked to configure them sequentially (in that case the user should change the failover timeout accordingly).	2013-11-18 11:30:08 +01:00
antirez	3a56013acb	Sentinel: state machine and timeouts simplified.	2013-11-18 11:12:58 +01:00
antirez	4be53b1c5d	Sentinel: election timeout define.	2013-11-18 10:08:06 +01:00
antirez	69d826a354	Sentinel: fix address of master in Hello messages. Once we switched configuration during a failover, we should advertise the new address. This was a serious race condition as the Sentinel performing the failover for a moment advertised the old address with the new configuration epoch: once trasmitted to the other Sentinels the broken configuration would remain there forever, until the next failover (because a greater configuration epoch is required to overwrite an older one).	2013-11-14 10:25:55 +01:00
antirez	e4c65e72c6	Sentinel: master address selection in get-master-address refactored.	2013-11-14 10:23:54 +01:00
antirez	c0d7229364	Sentinel: fix conditional to only affect slaves with wrong master.	2013-11-14 10:23:05 +01:00
antirez	dfbd9c5aeb	Sentinel: simplify and refactor slave reconfig code.	2013-11-14 00:36:43 +01:00
antirez	64ad6648a8	Sentinel: reconfigure slaves to right master.	2013-11-14 00:29:38 +01:00
antirez	3e27d678da	Sentinel: remember last time slave changed master.	2013-11-14 00:20:15 +01:00
antirez	8297745fa6	Sentinel: redirect-to-master is not ok with new algorithm. Now Sentinel believe the current configuration is always the winner and should be applied by Sentinels instead of trying to adapt our view of the cluster based on what we observe. So the only way to modify what a Sentinel believe to be the truth is to win an election and advertise the new configuration via Pub / Sub with a greater configuration epoch.	2013-11-13 17:03:48 +01:00
antirez	76a88f56e5	Sentinel: safer slave reconfig, master reported role should match.	2013-11-13 17:02:09 +01:00
antirez	ddaad9fe2d	Sentinel: role reporting fixed and added in SENTINEL output.	2013-11-13 16:39:57 +01:00
antirez	a0afa66f4b	Sentinel: being a master and reporting as slave is considered SDOWN.	2013-11-13 16:28:56 +01:00
antirez	17718fdcba	Sentinel: make sure role_reported is always updated.	2013-11-13 16:21:58 +01:00
antirez	46a053d34b	Sentinel: track role change time. Wait before reconfigurations.	2013-11-13 16:18:23 +01:00
antirez	9e40c46f5e	Sentinel: fix no-down check in master->slave conversion code.	2013-11-13 13:43:59 +01:00
antirez	ae35b7e240	Sentinel: readd slaves back after a master reset.	2013-11-13 13:01:11 +01:00
antirez	6bd4f6bffe	Sentinel: sentinelResetMaster() new flag to avoid removing set of sentinels. This commit also removes some dead code and cleanup generic flags.	2013-11-13 10:30:45 +01:00
antirez	1569af1f23	Sentinel: receive Pub/Sub messages from slaves.	2013-11-12 23:07:33 +01:00
antirez	dfa5f8b777	Sentinel: change event name when converting master to slave.	2013-11-12 23:00:17 +01:00
antirez	24158d1488	Sentinel: added config-epoch to SENTINEL masters output.	2013-11-12 17:22:04 +01:00
antirez	d2bc6dc39a	Sentinel: new failover algo, desync slaves and update config epoch.	2013-11-12 17:07:31 +01:00
antirez	4a128b949d	Sentinel: when starting failover seek for votes ASAP.	2013-11-12 16:38:02 +01:00
antirez	e6b9d5e97e	Sentinel: +new-epoch events.	2013-11-12 13:35:25 +01:00
antirez	54c447be52	Sentinel: wait some time between failover attempts.	2013-11-12 13:30:31 +01:00
antirez	ab4b2ec88f	Sentinel: allow to vote for myself.	2013-11-12 11:32:40 +01:00
antirez	b6b65b29c0	Sentinel: fix PUBLISH to masters and slaves.	2013-11-12 11:12:48 +01:00
antirez	90ab62fd5e	Sentinel: epoch introduced in leader vote.	2013-11-12 11:09:35 +01:00
antirez	8c1bf9a2bd	Sentinel: leadership handling changes WIP. Changes to leadership handling. Now the leader gets selected by every Sentinel, for a specified epoch, when the SENTINEL is-master-down-by-addr is sent. This command now includes the runid and the currentEpoch of the instance seeking for a vote. The Sentinel only votes a single time in a given epoch. Still a work in progress, does not even compile at this stage.	2013-11-11 18:30:14 +01:00
antirez	0bac36d0a1	Sentinel: handle Hello messages received via slaves correctly. Even when messages are received via the slave, we should perform operations (like adding a new Sentinel) in the context of the master.	2013-11-11 17:12:27 +01:00
antirez	9e1b27d49e	Sentinel: remove code not useful in the new design.	2013-11-11 12:06:11 +01:00
antirez	b93b0adc89	Sentinel: epoch introduced. Sentinel state now includes the idea of current epoch and config epoch. In the Hello message, that is now published both on masters and slaves, a Sentinel no longer just advertises itself but also broadcasts its current view of the configuration: the master name / ip / port and its current epoch. Sentinels receiving such information switch to the new master if the configuration epoch received is newer and the ip / port of the master are indeed different compared to the previos ones.	2013-11-11 11:05:58 +01:00
antirez	80da056c29	Sentinel: sentinelSendSlaveOf() was missing a var and the prototype.	2013-11-06 11:23:53 +01:00
antirez	23800d9e49	Sentinel: increment pending_commands counter in two more places. AUTH and SCRIPT KILL were sent without incrementing the pending commands counter. Clearly this needs some kind of wrapper doing it for the caller in order to be less bug prone.	2013-11-06 11:21:44 +01:00
antirez	671c1dfb56	Sentinel: always send CONFIG REWRITE when changing instance role. This change makes Sentinel less fragile about a number of failure modes. This commit also fixes a different bug as a side effect, SLAVEOF command was sent multiple times without incrementing the pending commands count.	2013-11-06 11:13:27 +01:00
antirez	fb9b76fe14	Cluster: slave node now uses the new protocol to get elected.	2013-09-26 11:13:17 +02:00
antirez	6ea8e0949c	sdsrange() does not need to return a value. Actaully the string is modified in-place and a reallocation is never needed, so there is no need to return the new sds string pointer as return value of the function, that is now just "void".	2013-07-24 11:21:39 +02:00
antirez	73ae8558c1	Sentinel: embed IPv6 address into [] when naming slave/sentinel instance.	2013-07-11 16:38:40 +02:00
antirez	3fc7f324d2	Sentinel: use comma as separator to publish hello messages. We use comma to play well with IPv6 addresses, but the implementation is still able to parse the old messages separated by colons.	2013-07-11 16:37:47 +02:00
antirez	5c5ebb0b9a	Sentinel: make sure published addr/id buffer is large enough. With ipv6 support we need more space, so we account for the IP address max size plus what we need for the Run ID, port, flags.	2013-07-10 14:44:38 +02:00
antirez	631d656a94	All IP string repr buffers are now REDIS_IP_STR_LEN bytes.	2013-07-09 11:32:52 +02:00
Geoff Garside	e04fdf26fe	Add IPv6 support to sentinel.c. This has been done by exposing the anetSockName() function anet.c to be used when the sentinel is publishing its existence to the masters. This implementation is very unintelligent as it will likely break if used with IPv6 as the nested colons will break any parsing of the PUBLISH string by the master.	2013-07-08 16:08:36 +02:00
Geoff Garside	2345cee335	Update calls to anetResolve to include buffer size	2013-07-08 15:57:22 +02:00
antirez	4c0f8c4e5a	Sentinel: parse new INFO replication output correctly. Sentinel was not able to detect slaves when connected to a very recent version of Redis master since a previos non-backward compatible change to INFO broken the parsing of the slaves ip:port INFO output. This fixes issue #1164	2013-06-20 10:23:23 +02:00
antirez	e5ef85c444	Sentinel: changes to tilt mode. Tilt mode was too aggressive (not processing INFO output), this resulted in a few problems: 1) Redirections were not followed when in tilt mode. This opened a window to misinform clients about the current master when a Sentinel was in tilt mode and a fail over happened during the time it was not able to update the state. 2) It was possible for a Sentinel exiting tilt mode to detect a false fail over start, if a slave rebooted with a wrong configuration about at the same time. This used to happen since in tilt mode we lose the information that the runid changed (reboot). Now instead the Sentinel in tilt mode will still remove the instance from the list of slaves if it changes state AND runid at the same time. Both are edge conditions but the changes should overall improve the reliability of Sentinel.	2013-04-30 15:08:29 +02:00
antirez	ef05a78e7e	Sentinel: more sensible delay in master demote after tilt.	2013-04-30 15:08:22 +02:00
antirez	48ede0d84d	Sentinel: only demote old master into slave under certain conditions. We used to always turn a master into a slave if the DEMOTE flag was set, as this was a resurrecting master instance. However the following race condition is possible for a Sentinel that got partitioned or internal issues (tilt mode), and was not able to refresh the state in the meantime: 1) Sentinel X is running, master is instance "A". 3) "A" fails, sentinels will promote slave "B" as master. 2) Sentinel X goes down because of a network partition. 4) "A" returns available, Sentinels will demote it as a slave. 5) "B" fails, other Sentinels will promote slave "A" as master. 6) At this point Sentinel X comes back. When "X" comes back he thinks that: "B" is the master. "A" is the slave to demote. We want to avoid that Sentinel "X" will demote "A" into a slave. We also want that Sentinel "X" will detect that the conditions changed and will reconfigure itself to monitor the right master. There are two main ways for the Sentinel to reconfigure itself after this event: 1) If "B" is reachable and already configured as a slave by other sentinels, "X" will perform a redirection to "A". 2) If there are not the conditions to demote "A", the fact that "A" reports to be a master will trigger a failover detection in "X", that will end into a reconfiguraiton to monitor "A". However if the Sentinel was not reachable, its state may not be updated, so in case it titled, or was partiitoned from the master instance of the slave to demote, the new implementation waits some time (enough to guarantee we can detect the new INFO, and new DOWN conditions). If after some time still there are not the right condiitons to demote the instance, the DEMOTE flag is cleared.	2013-04-26 17:02:13 +02:00
antirez	1965e22aa1	Sentinel: always redirect on master->slave transition. Sentinel redirected to the master if the instance changed runid or it was the first time we got INFO, and a role change was detected from master to slave. While this is a good idea in case of slave->master, since otherwise we could detect a failover without good reasons just after a reboot with a slave with a wrong configuration, in the case of master->slave transition is much better to always perform the redirection for the following reasons: 1) A Sentinel may go down for some time. When it is back online there is no other way to understand there was a failover. 2) Pointing clients to a slave seems to be always the wrong thing to do. 3) There is no good rationale about handling things differently once an instance is rebooted (runid change) in that case.	2013-04-24 11:30:17 +02:00
antirez	8e222c888f	Sentinel: turn old master into a slave when it comes back.	2013-04-19 16:47:24 +02:00
antirez	089cbe643f	Sentinel: advertise the promoted slave address only after successful setup.	2013-01-31 17:19:21 +01:00
guiquanz	9d09ce3981	Fixed many typos.	2013-01-19 10:59:44 +01:00
antirez	4365e5b2d3	BSD license added to every C source and header file.	2012-11-08 18:31:32 +01:00
antirez	db100c4671	Sentinel: Support for AUTH.	2012-09-26 18:59:54 +02:00
antirez	9bd0e097aa	Sentinel: reply -IDONTKNOW to get-master-addr-by-name on lack of info. If we don't have any clue about a master since it never replied to INFO so far, reply with an -IDONTKNOW error to SENTINEL get-master-addr-by-name requests.	2012-09-04 16:06:53 +02:00
antirez	8bdde086ac	Sentinel: more easy master redirection if master is a slave. Before this commit Sentienl used to redirect master ip/addr if the current instance reported to be a slave only if this was the first INFO output received, and the role was found to be slave. Now instead also if we find that the runid is different, and the reported role is slave, we also redirect to the reported master ip/addr. This unifies the behavior of Sentinel in the case of a reboot (where it will see the first INFO output with the wrong role and will perform the redirection), with the behavior of Sentinel in the case of a change in what it sees in the INFO output of the master.	2012-09-04 15:52:04 +02:00
antirez	6276434ad2	Sentinel: do not crash against slaves not publishing the runid. Older versions of Redis (before 2.4.17) don't publish the runid field in INFO. This commit makes Sentinel able to handle that without crashing.	2012-08-30 18:01:52 +02:00
antirez	58186b9dcf	Sentinel: INFO command implementation.	2012-08-29 12:44:24 +02:00
antirez	3ec701e059	Sentinel: Sentinel-side support for slave priority. The slave priority that is now published by Redis in INFO output is now used by Sentinel in order to select the slave with minimum priority for promotion, and in order to consider slaves with priority set to 0 as not able to play the role of master (they will never be promoted by Sentinel). The "slave-priority" field is now one of the fileds that Sentinel publishes when describing an instance via the SENTINEL commands such as "SENTINEL slaves mastername".	2012-08-28 17:45:01 +02:00
antirez	c14e0ecafd	Sentinel: suppress harmless warning by initializing 'table' to NULL. Note that the assertion guarantees that one of the if branches setting table is always entered.	2012-08-28 12:56:05 +02:00
antirez	850789ce73	Sentinel: send SCRIPT KILL on -BUSY reply and SDOWN instance. From the point of view of Redis an instance replying -BUSY is down, since it is effectively not able to reply to user requests. However a looping script is a recoverable condition in Redis if the script still did not performed any write to the dataset. In that case performing a fail over is not optimal, so Sentinel now tries to restore the normal server condition killing the script with a SCRIPT KILL command. If the script already performed some write before entering an infinite (or long enough to timeout) loop, SCRIPT KILL will not work and the fail over will be triggered anyway.	2012-08-24 12:29:54 +02:00
antirez	01477753e6	Sentinel: fixed a crash on script execution. The call to sentinelScheduleScriptExecution() lacked the final NULL argument to signal the end of arguments. This resulted into a crash.	2012-08-24 12:10:24 +02:00
antirez	cada7f9671	Sentinel: SENTINEL FAILOVER command implemented. This command can be used in order to force a Sentinel instance to start a failover for the specified master, as leader, forcing the failover even if the master is up. The commit also adds some minor refactoring and other improvements to functions already implemented that make them able to work when the master is not in SDOWN condition. For instance slave selection assumed that we ask INFO every second to every slave, this is true only when the master is in SDOWN condition, so slave selection did not worked when the master was not in SDOWN condition.	2012-08-03 12:41:27 +02:00
antirez	6275004ca6	Sentinel: client reconfiguration script execution. This commit adds support to optionally execute a script when one of the following events happen: * The failover starts (with a slave already promoted). * The failover ends. * The failover is aborted. The script is called with enough parameters (documented in the example sentinel.conf file) to provide information about the old and new ip:port pair of the master, the role of the sentinel (leader or observer) and the name of the master. The goal of the script is to inform clients of the configuration change in a way specific to the environment Sentinel is running, that can't be implemented in a genereal way inside Sentinel itself.	2012-08-02 18:40:30 +02:00
antirez	fd92b366b0	Sentinel: when leader in wait-start, sense another leader as race. When we are in wait start, if another leader (or any other external entity) turns a slave into a master, abort the failover, and detect it as an observer. Note that the wait-start state is mainly there for this reason but the abort was yet not implemented. This adds a new sentinel event -failover-abort-race.	2012-07-31 17:11:26 +02:00
antirez	91c15ed1b5	Sentinel: sentinelRefreshInstanceInfo() comments improved a bit.	2012-07-31 16:18:15 +02:00
antirez	75084e057d	Sentinel: abort failover when in wait-start if master is back. When we are a Leader Sentinel in wait-start state, starting with this commit the failover is aborted if the master returns online. This improves the way we handle a notable case of net split, that is the split between Sentinels and Redis servers, that will be a very common case of split becase Sentinels will often be installed in the client's network and servers can be in a differnt arm of the network. When Sentinels and Redis servers are isolated the master is in ODOWN condition since the Sentinels can agree about this state, however the failover does not start since there are no good slaves to promote (in this specific case all the slaves are unreachable). However when the split is resolved, Sentinels may sense the slave back a moment before they sense the master is back, so the failover may start without a good reason (since the master is actually working too). Now this condition is reversible, so the failover will be aborted immediately after if the master is detected to be working again, that is, not in SDOWN nor in ODOWN condition.	2012-07-31 10:19:34 +02:00
antirez	7f5bdba434	Merge remote-tracking branch 'origin/unstable' into unstable	2012-07-28 20:55:17 +02:00
antirez	3f194a9d25	Sentinel: scripts execution engine improved. We no longer use a vanilla fork+execve but take a queue of jobs of scripts to execute, with retry on error, timeouts, and so forth. Currently this is used only for notifications but soon the ability to also call clients reconfiguration scripts will be added.	2012-07-28 20:54:27 +02:00
Jan-Erik Rediger	c6c19c8372	Include sys/wait.h to avoid compiler warning gcc warned about an implicit declaration of function 'wait3'. Including this header fixes this.	2012-07-28 12:33:01 +03:00
antirez	ce7b838fb9	Sentinel: don't start a failover as leader if there is no good slave.	2012-07-26 12:09:40 +02:00
antirez	baace5fc42	Sentinel: ability to execute notification scripts.	2012-07-25 16:33:37 +02:00
antirez	672102c2ce	Sentinel: abort failover if no good slave is available. The previous behavior of the state machine was to wait some time and retry the slave selection, but this is not robust enough against drastic changes in the conditions of the monitored instances. What we do now when the slave selection fails is to abort the failover and return back monitoring the master. If the ODOWN condition is still present a new failover will be triggered and so forth. This commit also refactors the code we use to abort a failover.	2012-07-25 11:32:19 +02:00
antirez	9e5bef38e6	Sentinel: reset pending_commands in a more generic way.	2012-07-24 18:57:26 +02:00
antirez	a23a5b6c7d	Prevent a spurious +sdown event on switch. When we reset the master we should start with clean timestamps for ping replies otherwise we'll detect a spurious +sdown event, because on +master-switch event the previous master instance was probably in +sdown condition. Since we updated the address we should count time from scratch again. Also this commit makes sure to explicitly reset the count of pending commands, now we can do this because of the new way the hiredis link is closed.	2012-07-24 18:46:04 +02:00
antirez	d918e6f127	Sentinel: debugging message removed.	2012-07-24 18:20:05 +02:00
antirez	75fb6e5b8a	Sentinel: changes to connection handling and redirection. We disconnect the Redis instances hiredis link in a more robust way now. Also we change the way we perform the redirection for the +switch-master event, that is not just an instance reset with an address change. Using the same system we now implement the +redirect-to-master event that is triggered by an instance that is configured to be master but found to be a slave at the first INFO reply. In that case we monitor the master instead, logging the incident as an event.	2012-07-24 18:15:44 +02:00
antirez	2179c26916	Sentinel: check that instance still exists in reply callbacks. We can't be sure the instance object still exists when the reply callback is called.	2012-07-24 16:37:57 +02:00
antirez	d876d6feac	Sentinel: more robust failover detection as observer. Sentinel observers detect failover checking if a slave attached to the monitored master turns into its replication state from slave to master. However while this change may in theory only happen after a SLAVEOF NO ONE command, in practie it is very easy to reboot a slave instance with a wrong configuration that turns it into a master, especially if it was a past master before a successfull failover. This commit changes the detection policy so that if an instance goes from slave to master, but at the same time the runid has changed, we sense a reboot, and in that case we don't detect a failover at all. This commit also introduces the "reboot" sentinel event, that is logged at "warning" level (so this will trigger an admin notification). The commit also fixes a problem in the disconnect handler that assumed that the instance object always existed, that is not the case. Now we no longer assume that redisAsyncFree() will call the disconnection handler before returning.	2012-07-24 12:42:40 +02:00
antirez	6b5daa2df2	First implementation of Redis Sentinel. This commit implements the first, beta quality implementation of Redis Sentinel, a distributed monitoring system for Redis with notification and automatic failover capabilities. More info at http://redis.io/topics/sentinel	2012-07-23 13:14:44 +02:00

1 2 3 4 5

238 Commits