0
我遇到了ipython集群的怪异行为。计算结束,但许多结果永远不会到达客户端(并且在完成第一次计算后,引擎只是闲置)。ipython 0.13 zmq错误
我怀疑东西是错误的,因为ZMQ 1)不时我看到了以下错误:
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/parallel/client/asyncresult.py", line 118, in get
if not self.ready():
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/parallel/client/asyncresult.py", line 132, in ready
self.wait(0)
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/parallel/client/asyncresult.py", line 142, in wait
self._ready = self._client.wait(self.msg_ids, timeout)
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/parallel/client/client.py", line 1058, in wait
self.spin()
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/parallel/client/client.py", line 1015, in spin
self._flush_results(self._task_socket)
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/parallel/client/client.py", line 814, in _flush_results
idents,msg = self.session.recv(sock, mode=zmq.NOBLOCK)
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/zmq/session.py", line 642, in recv
idents, msg_list = self.feed_identities(msg_list, copy)
File "/data/misc/nano/python/env_stable/lib/python2.7/site-packages/IPython/zmq/session.py", line 673, in feed_identities
idx = msg_list.index(DELIM)
ValueError: '<IDS|MSG>' is not in list
Additionally IPython.zmq has two test failures:
======================================================================
ERROR: test_send (IPython.zmq.tests.test_session.TestSession)
----------------------------------------------------------------------
Traceback (most recent call last):
File "/clusterdata/python/env_stable/lib/python2.7/site-packages/IPython/zmq/tests/test_session.py", line 76, in test_send
socket = MockSocket(zmq.Context.instance(),zmq.PAIR)
File "/clusterdata/python/env_stable/lib/python2.7/site-packages/IPython/zmq/tests/test_session.py", line 34, in __init__
self.data = []
File "/clusterdata/python/env_stable/lib/python2.7/site-packages/zmq/sugar/attrsettr.py", line 38, in __setattr__
self.__class__.__name__, upper_key)
AttributeError: MockSocket has no such option: DATA
======================================================================
ERROR: test_send (IPython.zmq.tests.test_session.TestSession)
----------------------------------------------------------------------
Traceback (most recent call last):
File "/clusterdata/python/env_stable/lib/python2.7/site-packages/zmq/tests/__init__.py", line 108, in tearDown
raise RuntimeError("context could not terminate, open sockets likely remain in test")
RuntimeError: context could not terminate, open sockets likely remain in test
----------------------------------------------------------------------
我用pyzmq 13.0.0(如安装由PIP),以及zeromq 3.2.2 ,由pyzmq的设置编译。我使用ipython 13.1和python 2.7.3。
这是什么可能的任何建议,如果不是我怎么能找出更多的信息为什么会发生这些错误?
更新:事实证明,减速是由于ipcontroller的长任务队列,然后采取100%的CPU和滞后可怕。这是一个单独的问题,但我仍然会对上述的反馈感到满意。
MockSocket错误只影响测试本身,并在0.13.2 [此处发布候选版本](http://archive.ipython.org/testing/0.13.2)中修复。 – minrk 2013-03-27 21:34:53
不知道其他错误可能是什么?另外,根据更新,ipcontroller是否应该像地狱一样滞留4000个职位(如果它不会滞后几百个)? – 2013-03-27 22:09:32
这显然不应该,但这并不意味着你的系统出了问题。如果你有很多工作,我强烈建议将TaskScheduler.hwm设置为大。 – minrk 2013-03-28 05:09:13