Compare commits
52
Commits
74ebb6e250
..
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c6abf8e98d | ||
|
|
6095ca6fec | ||
|
|
db6cdeed21 | ||
|
|
e0f4aa3c13 | ||
|
|
d8e7c4c80f | ||
|
|
048ba027f3 | ||
|
|
35864d4b38 | ||
|
|
f421c92f83 | ||
|
|
164c72504c | ||
|
|
3bb3ffa3a0 | ||
|
|
5afa8b22b2 | ||
|
|
78ac536bb9 | ||
|
|
2032d2e91d | ||
|
|
292b12c5a1 | ||
|
|
ede0f80cea | ||
|
|
1bf587c4da | ||
|
|
6d465c69c9 | ||
|
|
8132bb9321 | ||
|
|
bde7869f97 | ||
|
|
7302d59727 | ||
|
|
f02163d88c | ||
|
|
06bfcfe065 | ||
|
|
d8856e7dc9 | ||
|
|
4fede26967 | ||
|
|
1ef66a7674 | ||
|
|
555a4780e3 | ||
|
|
0fe5c35dda | ||
|
|
e5ac7e0b05 | ||
|
|
fcd502d05b | ||
|
|
df8b2d836e | ||
|
|
910acc2a15 | ||
|
|
71809b7374 | ||
|
|
d73c7798b3 | ||
|
|
89e5f1369e | ||
|
|
ba796a1d6e | ||
|
|
e371d871e6 | ||
|
|
70536b79e7 | ||
|
|
66c9509c27 | ||
|
|
e09a588e77 | ||
|
|
61160c8595 | ||
|
|
83eb759b9c | ||
|
|
811b2bb1e3 | ||
|
|
da655346cd | ||
|
|
e8c3b5dac0 | ||
|
|
d4a1a35190 | ||
|
|
678d6f620a | ||
|
|
29f96b1e8c | ||
|
|
6636d8e1e9 | ||
|
|
100e82bd0f | ||
|
|
a02878c5ab | ||
|
|
fd69dd4a5f | ||
|
|
0412afb69c |
Binary file not shown.
@@ -0,0 +1,21 @@
|
|||||||
|
MIT License
|
||||||
|
|
||||||
|
Copyright (c) 2022 Dominik Roth
|
||||||
|
|
||||||
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||||
|
of this software and associated documentation files (the "Software"), to deal
|
||||||
|
in the Software without restriction, including without limitation the rights
|
||||||
|
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||||
|
copies of the Software, and to permit persons to whom the Software is
|
||||||
|
furnished to do so, subject to the following conditions:
|
||||||
|
|
||||||
|
The above copyright notice and this permission notice shall be included in all
|
||||||
|
copies or substantial portions of the Software.
|
||||||
|
|
||||||
|
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||||
|
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||||
|
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||||
|
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||||
|
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||||
|
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||||
|
SOFTWARE.
|
||||||
@@ -14,6 +14,12 @@ Project Columbus is a framework for trivial 2D OpenAI Gym environments that are
|
|||||||
pip install -e .
|
pip install -e .
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
Columbus.pdf contains a overview of columbus.
|
||||||
|
|
||||||
|
## Layout of the Repo
|
||||||
|
|
||||||
### env.py
|
### env.py
|
||||||
|
|
||||||

|

|
||||||
@@ -21,23 +27,22 @@ Contains the ColumbusEnv.
|
|||||||
There exist two ways to implement new envs:
|
There exist two ways to implement new envs:
|
||||||
|
|
||||||
- Subclassing ColumbusEnv and expanding _init_ and overriding _setup_.
|
- Subclassing ColumbusEnv and expanding _init_ and overriding _setup_.
|
||||||
- Using the ColumbusConfigDefined with a desired configuration. This makes configuring ColumbusEnvs via ClusterWorks2-configs possible. (See ColumbusConfigDefinedExample.md for an example of how the parameters are supposed to look like (uses yaml format), I don't have to to write a better documentation right now...)
|
- Using the ColumbusConfigDefined with a desired configuration. This makes configuring ColumbusEnvs via ClusterWorks2-configs possible. (See configs/example.yaml for an example of how the parameters are supposed to look like (uses yaml format)
|
||||||
|
- We now support using units (px, em, ct) in config files, examples can be found in configs/Example_Units.yaml
|
||||||
|
- The environments used in my thesis can also be found in configs/
|
||||||
|
|
||||||
##### Some caveats / infos
|
##### Some caveats / infos
|
||||||
|
|
||||||
- If you want to render to a window (pygame-gui) call render with mode='human'
|
- If you want to render to a window (pygame-gui) call render with mode='human'
|
||||||
- If you want visualize the covariance you have supply the cholesky-decomp of the cov-matrix to render
|
- If you want visualize the covariance you have to supply the cholesky-decomp of the cov-matrix to render
|
||||||
- If you want to render into a mp4, you have to call render with a mode!='human' and assemble/encode the returned frames yourself into a mp4/webm/...
|
- If you want to render into a mp4, you have to call render with a mode!='human' and assemble/encode the returned frames yourself into a mp4/webm/...
|
||||||
- Even while the agent plays, some keyboard-inputs are possible (to test the agents reaction to situations he would never enter by itself. Look at \_handle_user_input in env.py for avaible keys)
|
- Even while the agent plays, some keyboard-inputs are possible (to test the agents reaction to situations he would never enter by itself. Look at \_handle_user_input in env.py for avaible keys)
|
||||||
|
- The sampling-rate of the physics engine is bound to the frame-rate of the rendering engine (1:1). This means too low fps / too fast agents / too thin barriers will lead to the agent tunneling through barriers. You can fix this by setting a higher agent-drag (which decreases the maximum speed) or making barriers thicker. A feature allowing the physics engine to sample multiple smaller steps within a single rendering step could be added in the future.
|
||||||
|
|
||||||
### entities.py
|
### entities.py
|
||||||
|
|
||||||
Contains all implemented entities (e.g. the Agent, Rewards and Enemies)
|
Contains all implemented entities (e.g. the Agent, Rewards and Enemies)
|
||||||
|
|
||||||
##### Some caveats
|
|
||||||
|
|
||||||
- Support for non spherical entities (rectangles) is very new. There might be bugs that I have not yet found
|
|
||||||
|
|
||||||
### observables.py
|
### observables.py
|
||||||
|
|
||||||
Contains all 'oberservables'. These are attached to envs to define what kind of output is given to the agent. This way environments can be designed independently from the observation machanism that is used by the agent to play it.
|
Contains all 'oberservables'. These are attached to envs to define what kind of output is given to the agent. This way environments can be designed independently from the observation machanism that is used by the agent to play it.
|
||||||
@@ -45,11 +50,8 @@ Contains all 'oberservables'. These are attached to envs to define what kind of
|
|||||||
##### Some caveats
|
##### Some caveats
|
||||||
|
|
||||||
- CNNObservable seems to be broken currently. (Fixing it is also no priority for me)
|
- CNNObservable seems to be broken currently. (Fixing it is also no priority for me)
|
||||||
|
- RayObservable is using a naive ray-marching (basicaly just line-sweeping). For large amounts of rays this turn out to be the computational bottleneck of the environment. Switching to a more efficient algorithm (based on euclidean formulars and line intersects) would be possible in the future...
|
||||||
|
|
||||||
### humanPlayer.py
|
### humanPlayer.py
|
||||||
|
|
||||||
Allows environments to be played by a human using mouse input.
|
Allows environments to be played by a human using mouse input. Now even works for ColumbusConfigDefined.
|
||||||
|
|
||||||
##### Some caveats
|
|
||||||
|
|
||||||
- Does not yet work for ColumbusConfigDefined...
|
|
||||||
|
|||||||
+250
-11
@@ -3,30 +3,53 @@ import math
|
|||||||
|
|
||||||
|
|
||||||
class Entity(object):
|
class Entity(object):
|
||||||
|
def __call__(cls, *args, **kwargs):
|
||||||
|
obj = type.__call__(cls, *args, **kwargs)
|
||||||
|
obj.__post_init__()
|
||||||
|
return obj
|
||||||
|
|
||||||
def __init__(self, env):
|
def __init__(self, env):
|
||||||
self.shape = None
|
self.shape = None
|
||||||
self.env = env
|
self.env = env
|
||||||
self.pos = (env.random(), env.random())
|
self.pos = (env.random(), env.random())
|
||||||
|
self.last_pos = None
|
||||||
self.speed = (0, 0)
|
self.speed = (0, 0)
|
||||||
self.acc = (0, 0)
|
self.acc = (0, 0)
|
||||||
self.drag = 0
|
self.drag = 0
|
||||||
self.col = (255, 255, 255)
|
self.col = (255, 255, 255)
|
||||||
self.solid = False
|
self.solid = False
|
||||||
self.movable = False # False = Non movable, True = Movable, x>1: lighter movable
|
self.movable = False # False = Non movable, True = Movable, x>1: lighter movable
|
||||||
|
self.void_collidable = False
|
||||||
self.elasticity = 1
|
self.elasticity = 1
|
||||||
#self.collision_changes_speed = True
|
|
||||||
self.collision_changes_speed = self.env.controll_type == 'ACC'
|
self.collision_changes_speed = self.env.controll_type == 'ACC'
|
||||||
self.collision_elasticity = self.env.default_collision_elasticity
|
self.collision_elasticity = self.env.default_collision_elasticity
|
||||||
self._crash_list = []
|
self._crash_list = []
|
||||||
self._coll_add_pushback = 0
|
self._coll_add_pushback = 0
|
||||||
|
self.crash_conservation_of_energy = True
|
||||||
|
self.draw_path = False
|
||||||
|
self.draw_path_col = [int(c/5) for c in self.col]
|
||||||
|
self.draw_path_width = 2
|
||||||
|
self.draw_path_harm = False
|
||||||
|
self.draw_path_harm_col = [c for c in self.draw_path_col]
|
||||||
|
self.draw_path_harm_col[0] += int(255/3)
|
||||||
|
self.min_speed = 0
|
||||||
|
self.max_speed = math.inf
|
||||||
|
|
||||||
|
def __post_init__(self):
|
||||||
|
pass
|
||||||
|
|
||||||
def physics_step(self):
|
def physics_step(self):
|
||||||
x, y = self.pos
|
x, y = self.pos
|
||||||
vx, vy = self.speed
|
vx, vy = self.speed
|
||||||
ax, ay = self.acc
|
ax, ay = self.acc
|
||||||
vx, vy = vx+ax*self.env.acc_fac, vy+ay*self.env.acc_fac
|
vx, vy = vx+ax*self.env.acc_fac, vy+ay*self.env.acc_fac
|
||||||
|
speeds = math.sqrt(vx**2 + vy**2)
|
||||||
|
if speeds < self.min_speed:
|
||||||
|
vx, vy = vx/speeds*self.min_speed, vy/speeds*self.min_speed
|
||||||
|
if speeds > self.max_speed:
|
||||||
|
vx, vy = vx/speeds*self.max_speed, vy/speeds*self.max_speed
|
||||||
x, y = x+vx*self.env.speed_fac, y+vy*self.env.speed_fac
|
x, y = x+vx*self.env.speed_fac, y+vy*self.env.speed_fac
|
||||||
if not self.env.torus_topology:
|
if not self.env.torus_topology and self.void_collidable:
|
||||||
if x > 1 or x < 0:
|
if x > 1 or x < 0:
|
||||||
x, y, vx, vy = self.calc_void_collision(x < 0, x, y, vx, vy)
|
x, y, vx, vy = self.calc_void_collision(x < 0, x, y, vx, vy)
|
||||||
if y > 1 or y < 0:
|
if y > 1 or y < 0:
|
||||||
@@ -47,8 +70,19 @@ class Entity(object):
|
|||||||
self._crash_list = []
|
self._crash_list = []
|
||||||
|
|
||||||
def draw(self):
|
def draw(self):
|
||||||
raise Exception(
|
self._draw_path()
|
||||||
'[!] draw not implemented for shape "'+str(self.shape)+'"')
|
|
||||||
|
def _draw_path(self):
|
||||||
|
if self.draw_path and self.last_pos:
|
||||||
|
col = self.draw_path_col
|
||||||
|
if self.draw_path_harm:
|
||||||
|
if self.env.gotHarm:
|
||||||
|
col = self.draw_path_harm_col
|
||||||
|
pygame.draw.line(self.env.path_overlay, col,
|
||||||
|
(self.last_pos[0]*self.env.width, self.last_pos[1]*self.env.height), (self.pos[0]*self.env.width, self.pos[1]*self.env.height), self.draw_path_width)
|
||||||
|
pygame.draw.circle(self.env.path_overlay, col,
|
||||||
|
(self.pos[0]*self.env.width, self.pos[1]*self.env.height), max(0, self.draw_path_width/2-3))
|
||||||
|
self.last_pos = self.pos[0], self.pos[1]
|
||||||
|
|
||||||
def on_collision(self, other, depth):
|
def on_collision(self, other, depth):
|
||||||
if self.solid and other.solid:
|
if self.solid and other.solid:
|
||||||
@@ -66,10 +100,17 @@ class Entity(object):
|
|||||||
return
|
return
|
||||||
force_dir = force_dir[0]/force_dir_len, force_dir[1]/force_dir_len
|
force_dir = force_dir[0]/force_dir_len, force_dir[1]/force_dir_len
|
||||||
if not self.env.torus_topology:
|
if not self.env.torus_topology:
|
||||||
if self.env.agent.pos[0] > 0.99 or self.env.agent.pos[0] < 0.01:
|
if self == self.env.agent:
|
||||||
force_dir = force_dir[0], force_dir[1] * 2
|
agent = self
|
||||||
if self.env.agent.pos[1] > 0.99 or self.env.agent.pos[1] < 0.01:
|
elif other == self.env.agent:
|
||||||
force_dir = force_dir[0] * 2, force_dir[1]
|
agent = other
|
||||||
|
else:
|
||||||
|
agent = None
|
||||||
|
if agent:
|
||||||
|
if agent.pos[0] > 0.99 or agent.pos[0] < 0.01:
|
||||||
|
force_dir = force_dir[0], force_dir[1] * 2
|
||||||
|
if agent.pos[1] > 0.99 or agent.pos[1] < 0.01:
|
||||||
|
force_dir = force_dir[0] * 2, force_dir[1]
|
||||||
depth *= 1.0*self.movable/(self.movable + other.movable)/2
|
depth *= 1.0*self.movable/(self.movable + other.movable)/2
|
||||||
depth /= other.elasticity
|
depth /= other.elasticity
|
||||||
force_vec = force_dir[0]*depth/self.env.width, \
|
force_vec = force_dir[0]*depth/self.env.width, \
|
||||||
@@ -86,7 +127,7 @@ class Entity(object):
|
|||||||
force_vec[0]*self.collision_elasticity/self.env.speed_fac, self.speed[1] + \
|
force_vec[0]*self.collision_elasticity/self.env.speed_fac, self.speed[1] + \
|
||||||
force_vec[1]*self.collision_elasticity/self.env.speed_fac
|
force_vec[1]*self.collision_elasticity/self.env.speed_fac
|
||||||
newspeed = math.sqrt(self.speed[0]**2+self.speed[1]**2)
|
newspeed = math.sqrt(self.speed[0]**2+self.speed[1]**2)
|
||||||
if newspeed > oldspeed*1.1:
|
if self.crash_conservation_of_energy and newspeed > oldspeed*1.1:
|
||||||
self.speed = self.speed[0]/newspeed*1.1 * \
|
self.speed = self.speed[0]/newspeed*1.1 * \
|
||||||
oldspeed, self.speed[1]/newspeed*oldspeed*1.1
|
oldspeed, self.speed[1]/newspeed*oldspeed*1.1
|
||||||
|
|
||||||
@@ -113,6 +154,24 @@ class Entity(object):
|
|||||||
def kill(self):
|
def kill(self):
|
||||||
self.env.kill_entity(self)
|
self.env.kill_entity(self)
|
||||||
|
|
||||||
|
def getQuasiRadius(self):
|
||||||
|
raise Exception()
|
||||||
|
|
||||||
|
def getTop(self):
|
||||||
|
raise Exception()
|
||||||
|
|
||||||
|
def getBottom(self):
|
||||||
|
raise Exception()
|
||||||
|
|
||||||
|
def getLeft(self):
|
||||||
|
raise Exception()
|
||||||
|
|
||||||
|
def getRight(self):
|
||||||
|
raise Exception()
|
||||||
|
|
||||||
|
def getCenter(self):
|
||||||
|
raise Exception()
|
||||||
|
|
||||||
|
|
||||||
class CircularEntity(Entity):
|
class CircularEntity(Entity):
|
||||||
def __init__(self, env):
|
def __init__(self, env):
|
||||||
@@ -121,6 +180,7 @@ class CircularEntity(Entity):
|
|||||||
self.radius = 10
|
self.radius = 10
|
||||||
|
|
||||||
def draw(self):
|
def draw(self):
|
||||||
|
super().draw()
|
||||||
x, y = self.pos
|
x, y = self.pos
|
||||||
pygame.draw.circle(self.env.surface, self.col,
|
pygame.draw.circle(self.env.surface, self.col,
|
||||||
(x*self.env.width, y*self.env.height), self.radius, width=0)
|
(x*self.env.width, y*self.env.height), self.radius, width=0)
|
||||||
@@ -181,6 +241,24 @@ class CircularEntity(Entity):
|
|||||||
raise Exception(
|
raise Exception(
|
||||||
'[!] Shape "circle" does not know how to collide with shape "'+str(other.shape)+'"')
|
'[!] Shape "circle" does not know how to collide with shape "'+str(other.shape)+'"')
|
||||||
|
|
||||||
|
def getQuasiRadius(self):
|
||||||
|
return self.radius
|
||||||
|
|
||||||
|
def getTop(self):
|
||||||
|
return self.pos[1]*self.env.height - self.radius
|
||||||
|
|
||||||
|
def getBottom(self):
|
||||||
|
return self.pos[1]*self.env.height + self.radius
|
||||||
|
|
||||||
|
def getLeft(self):
|
||||||
|
return self.pos[0]*self.env.width - self.radius
|
||||||
|
|
||||||
|
def getRight(self):
|
||||||
|
return self.pos[0]*self.env.width + self.radius
|
||||||
|
|
||||||
|
def getCenter(self):
|
||||||
|
return self.pos[0]*self.env.width, self.pos[1]*self.env.height
|
||||||
|
|
||||||
|
|
||||||
class RectangularEntity(Entity):
|
class RectangularEntity(Entity):
|
||||||
def __init__(self, env):
|
def __init__(self, env):
|
||||||
@@ -190,6 +268,7 @@ class RectangularEntity(Entity):
|
|||||||
self.height = 10
|
self.height = 10
|
||||||
|
|
||||||
def draw(self):
|
def draw(self):
|
||||||
|
super().draw()
|
||||||
x, y = self.pos
|
x, y = self.pos
|
||||||
rect = pygame.Rect(x*self.env.width, y *
|
rect = pygame.Rect(x*self.env.width, y *
|
||||||
self.env.width, self.width, self.height)
|
self.env.width, self.width, self.height)
|
||||||
@@ -200,6 +279,58 @@ class RectangularEntity(Entity):
|
|||||||
raise Exception(
|
raise Exception(
|
||||||
'[!] Collisions in this direction not implemented for shape "rectangle"')
|
'[!] Collisions in this direction not implemented for shape "rectangle"')
|
||||||
|
|
||||||
|
def physics_step(self):
|
||||||
|
x, y = self.pos
|
||||||
|
vx, vy = self.speed
|
||||||
|
ax, ay = self.acc
|
||||||
|
vx, vy = vx+ax*self.env.acc_fac, vy+ay*self.env.acc_fac
|
||||||
|
speeds = math.sqrt(vx**2 + vy**2)
|
||||||
|
if speeds < self.min_speed:
|
||||||
|
vx, vy = vx/speeds*self.min_speed, vy/speeds*self.min_speed
|
||||||
|
if speeds > self.max_speed:
|
||||||
|
vx, vy = vx/speeds*self.max_speed, vy/speeds*self.max_speed
|
||||||
|
x, y = x+vx*self.env.speed_fac, y+vy*self.env.speed_fac
|
||||||
|
if not self.env.torus_topology and self.void_collidable:
|
||||||
|
if x+(self.width/self.env.width) > 1 or x < 0:
|
||||||
|
if x < 0:
|
||||||
|
x, y, vx, vy = self.calc_void_collision(
|
||||||
|
x < 0, x, y, vx, vy)
|
||||||
|
else:
|
||||||
|
x, y, vx, vy = self.calc_void_collision(
|
||||||
|
x < 0, x+(self.width/self.env.width), y, vx, vy)
|
||||||
|
x -= (self.width/self.env.width)
|
||||||
|
if y+(self.height/self.env.height) > 1 or y < 0:
|
||||||
|
if y < 0:
|
||||||
|
x, y, vx, vy = self.calc_void_collision(
|
||||||
|
2 + (x < 0), x, y, vx, vy)
|
||||||
|
else:
|
||||||
|
x, y, vx, vy = self.calc_void_collision(
|
||||||
|
2 + (x < 0), x, y+(self.height/self.env.height), vx, vy)
|
||||||
|
y -= (self.height/self.env.height)
|
||||||
|
else:
|
||||||
|
x = x % 1
|
||||||
|
y = y % 1
|
||||||
|
self.speed = vx/(1+self.drag), vy/(1+self.drag)
|
||||||
|
self.pos = x, y
|
||||||
|
|
||||||
|
def getQuasiRadius(self):
|
||||||
|
return self.width + self.height
|
||||||
|
|
||||||
|
def getTop(self):
|
||||||
|
return self.pos[1]*self.env.height
|
||||||
|
|
||||||
|
def getBottom(self):
|
||||||
|
return self.pos[1]*self.env.height + self.height
|
||||||
|
|
||||||
|
def getLeft(self):
|
||||||
|
return self.pos[0]*self.env.width
|
||||||
|
|
||||||
|
def getRight(self):
|
||||||
|
return self.pos[0]*self.env.width*self.env.height + self.width
|
||||||
|
|
||||||
|
def getCenter(self):
|
||||||
|
return self.pos[0]*self.env.width+self.width/2, self.pos[1]*self.env.height+self.height/2
|
||||||
|
|
||||||
|
|
||||||
class Agent(CircularEntity):
|
class Agent(CircularEntity):
|
||||||
def __init__(self, env):
|
def __init__(self, env):
|
||||||
@@ -210,6 +341,31 @@ class Agent(CircularEntity):
|
|||||||
self.controll_type = self.env.controll_type
|
self.controll_type = self.env.controll_type
|
||||||
self.solid = True
|
self.solid = True
|
||||||
self.movable = True
|
self.movable = True
|
||||||
|
self.void_collidable = True
|
||||||
|
|
||||||
|
def controll_step(self):
|
||||||
|
self._read_input()
|
||||||
|
self.env.check_collisions_for(self)
|
||||||
|
|
||||||
|
def _read_input(self):
|
||||||
|
if self.controll_type == 'SPEED':
|
||||||
|
self.speed = self.env.inp[0] - 0.5, self.env.inp[1] - 0.5
|
||||||
|
elif self.controll_type == 'ACC':
|
||||||
|
self.acc = self.env.inp[0] - 0.5, self.env.inp[1] - 0.5
|
||||||
|
else:
|
||||||
|
raise Exception('Unsupported controll_type')
|
||||||
|
|
||||||
|
|
||||||
|
# Does not work! Don't use!
|
||||||
|
class PongAgent(RectangularEntity):
|
||||||
|
def __init__(self, env):
|
||||||
|
super(PongAgent, self).__init__(env)
|
||||||
|
self.pos = (0.5, 0.5)
|
||||||
|
self.col = (0, 0, 255)
|
||||||
|
self.drag = self.env.agent_drag
|
||||||
|
self.controll_type = self.env.controll_type
|
||||||
|
self.solid = True
|
||||||
|
self.movable = True
|
||||||
|
|
||||||
def controll_step(self):
|
def controll_step(self):
|
||||||
self._read_input()
|
self._read_input()
|
||||||
@@ -217,9 +373,9 @@ class Agent(CircularEntity):
|
|||||||
|
|
||||||
def _read_input(self):
|
def _read_input(self):
|
||||||
if self.controll_type == 'SPEED':
|
if self.controll_type == 'SPEED':
|
||||||
self.speed = self.env.inp[0] - 0.5, self.env.inp[1] - 0.5
|
self.speed = 0, self.env.inp[1] - 0.5
|
||||||
elif self.controll_type == 'ACC':
|
elif self.controll_type == 'ACC':
|
||||||
self.acc = self.env.inp[0] - 0.5, self.env.inp[1] - 0.5
|
self.acc = 0, self.env.inp[1] - 0.5
|
||||||
else:
|
else:
|
||||||
raise Exception('Unsupported controll_type')
|
raise Exception('Unsupported controll_type')
|
||||||
|
|
||||||
@@ -321,6 +477,33 @@ class Collectable(CircularEntity):
|
|||||||
self.env.check_collisions_for(self)
|
self.env.check_collisions_for(self)
|
||||||
|
|
||||||
|
|
||||||
|
class RectCollectable(RectangularEntity):
|
||||||
|
def __init__(self, env):
|
||||||
|
super(RectCollectable, self).__init__(env)
|
||||||
|
self.avaible = True
|
||||||
|
self.enforce_not_on_barrier = False
|
||||||
|
self.reward = 10
|
||||||
|
self.collectors = []
|
||||||
|
|
||||||
|
def on_collision(self, other, depth):
|
||||||
|
super().on_collision(other, depth)
|
||||||
|
if isinstance(other, Barrier):
|
||||||
|
self.on_barrier_collision()
|
||||||
|
else:
|
||||||
|
for Col in self.collectors:
|
||||||
|
if isinstance(other, Col):
|
||||||
|
other.on_collect(self)
|
||||||
|
self.on_collected()
|
||||||
|
|
||||||
|
def on_collected(self):
|
||||||
|
self.env.new_reward += self.reward
|
||||||
|
|
||||||
|
def on_barrier_collision(self):
|
||||||
|
if self.enforce_not_on_barrier:
|
||||||
|
self.pos = (self.env.random(), self.env.random())
|
||||||
|
self.env.check_collisions_for(self)
|
||||||
|
|
||||||
|
|
||||||
class Reward(Collectable):
|
class Reward(Collectable):
|
||||||
def __init__(self, env):
|
def __init__(self, env):
|
||||||
super(Reward, self).__init__(env)
|
super(Reward, self).__init__(env)
|
||||||
@@ -335,6 +518,9 @@ class OnceReward(Reward):
|
|||||||
self.reward = 500
|
self.reward = 500
|
||||||
|
|
||||||
def on_collected(self):
|
def on_collected(self):
|
||||||
|
# Force rerender of value func (even in static envs)
|
||||||
|
self.env._invalidate_value_map()
|
||||||
|
|
||||||
self.env.new_abs_reward += self.reward
|
self.env.new_abs_reward += self.reward
|
||||||
self.kill()
|
self.kill()
|
||||||
|
|
||||||
@@ -346,11 +532,50 @@ class TeleportingReward(OnceReward):
|
|||||||
self.env.check_collisions_for(self)
|
self.env.check_collisions_for(self)
|
||||||
|
|
||||||
def on_collected(self):
|
def on_collected(self):
|
||||||
|
# Force rerender of value func (even in static envs)
|
||||||
|
self.env._invalidate_value_map()
|
||||||
|
|
||||||
self.env.new_abs_reward += self.reward
|
self.env.new_abs_reward += self.reward
|
||||||
self.pos = (self.env.random(), self.env.random())
|
self.pos = (self.env.random(), self.env.random())
|
||||||
self.env.check_collisions_for(self)
|
self.env.check_collisions_for(self)
|
||||||
|
|
||||||
|
|
||||||
|
class LoopReward(OnceReward):
|
||||||
|
def __init__(self, env):
|
||||||
|
super().__init__(env)
|
||||||
|
self.loop = [[0.25, 0.5], [0.75, 0.5]]
|
||||||
|
self.state = 0
|
||||||
|
self.jump_to_state()
|
||||||
|
self.barrier_physics = False
|
||||||
|
|
||||||
|
def jump_to_state(self):
|
||||||
|
# Force rerender of value func (even in static envs)
|
||||||
|
self.env._invalidate_value_map()
|
||||||
|
|
||||||
|
pos_vec = [v for v in self.loop[self.state]]
|
||||||
|
if len(pos_vec) == 4:
|
||||||
|
pos_vec = pos_vec[0] + pos_vec[2] * \
|
||||||
|
(self.env.random()-0.5), pos_vec[1] + \
|
||||||
|
pos_vec[3]*(self.env.random()-0.5)
|
||||||
|
self.pos = pos_vec
|
||||||
|
|
||||||
|
def next_state(self):
|
||||||
|
self.state = (self.state + 1) % len(self.loop)
|
||||||
|
|
||||||
|
def jump_next(self):
|
||||||
|
self.next_state()
|
||||||
|
self.jump_to_state()
|
||||||
|
|
||||||
|
def on_collected(self):
|
||||||
|
self.env.new_abs_reward += self.reward
|
||||||
|
self.jump_next()
|
||||||
|
|
||||||
|
def physics_step(self):
|
||||||
|
if self.barrier_physics:
|
||||||
|
self.env.check_collisions_for(self)
|
||||||
|
super().physics_step()
|
||||||
|
|
||||||
|
|
||||||
class TimeoutReward(OnceReward):
|
class TimeoutReward(OnceReward):
|
||||||
def __init__(self, env):
|
def __init__(self, env):
|
||||||
super(TimeoutReward, self).__init__(env)
|
super(TimeoutReward, self).__init__(env)
|
||||||
@@ -367,6 +592,9 @@ class TimeoutReward(OnceReward):
|
|||||||
|
|
||||||
def on_collected(self):
|
def on_collected(self):
|
||||||
if self.avaible:
|
if self.avaible:
|
||||||
|
# Force rerender of value func (even in static envs)
|
||||||
|
self.env._invalidate_value_map()
|
||||||
|
|
||||||
self.env.new_abs_reward += self.reward
|
self.env.new_abs_reward += self.reward
|
||||||
self.set_avaible(False)
|
self.set_avaible(False)
|
||||||
self.env.timers.append((self.timeout, self.set_avaible, True))
|
self.env.timers.append((self.timeout, self.set_avaible, True))
|
||||||
@@ -406,6 +634,14 @@ class Goal(Collectable):
|
|||||||
self.collectors = [Ball]
|
self.collectors = [Ball]
|
||||||
|
|
||||||
|
|
||||||
|
class RectGoal(RectCollectable):
|
||||||
|
def __init__(self, env):
|
||||||
|
super(RectGoal, self).__init__(env)
|
||||||
|
self.col = (0, 200, 0)
|
||||||
|
self.reward = 500
|
||||||
|
self.collectors = [Ball]
|
||||||
|
|
||||||
|
|
||||||
class TeleportingGoal(Goal):
|
class TeleportingGoal(Goal):
|
||||||
def __init__(self, env):
|
def __init__(self, env):
|
||||||
super(TeleportingGoal, self).__init__(env)
|
super(TeleportingGoal, self).__init__(env)
|
||||||
@@ -413,6 +649,9 @@ class TeleportingGoal(Goal):
|
|||||||
self.env.check_collisions_for(self)
|
self.env.check_collisions_for(self)
|
||||||
|
|
||||||
def on_collected(self):
|
def on_collected(self):
|
||||||
|
# Force rerender of value func (even in static envs)
|
||||||
|
self.env._invalidate_value_map()
|
||||||
|
|
||||||
self.env.new_abs_reward += self.reward
|
self.env.new_abs_reward += self.reward
|
||||||
self.pos = (self.env.random(), self.env.random())
|
self.pos = (self.env.random(), self.env.random())
|
||||||
self.env.check_collisions_for(self)
|
self.env.check_collisions_for(self)
|
||||||
|
|||||||
+313
-132
@@ -6,46 +6,17 @@ import pygame
|
|||||||
import random as random_dont_use
|
import random as random_dont_use
|
||||||
from os import urandom
|
from os import urandom
|
||||||
import math
|
import math
|
||||||
from columbus import entities, observables
|
|
||||||
import torch as th
|
import torch as th
|
||||||
|
|
||||||
|
from columbus import entities, observables
|
||||||
def parseObs(obsConf):
|
from columbus.utils import soft_int, parseObs
|
||||||
if type(obsConf) == list:
|
|
||||||
obs = []
|
|
||||||
for i, c in enumerate(obsConf):
|
|
||||||
obs.append(parseObs(c))
|
|
||||||
if len(obs) == 1:
|
|
||||||
return obs[0]
|
|
||||||
else:
|
|
||||||
return observables.CompositionalObservable(obs)
|
|
||||||
|
|
||||||
if obsConf['type'] == 'State':
|
|
||||||
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
|
||||||
return observables.StateObservable(**conf)
|
|
||||||
elif obsConf['type'] == 'Compass':
|
|
||||||
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
|
||||||
return observables.CompassObservable(**conf)
|
|
||||||
elif obsConf['type'] == 'RayCast':
|
|
||||||
chans = []
|
|
||||||
for chan in obsConf.get('chans', []):
|
|
||||||
chans.append(getattr(entities, chan))
|
|
||||||
conf = {k: v for k, v in obsConf.items() if k not in ['type', 'chans']}
|
|
||||||
return observables.RayObservable(chans=chans, **conf)
|
|
||||||
elif obsConf['type'] == 'CNN':
|
|
||||||
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
|
||||||
return observables.CnnObservable(**conf)
|
|
||||||
elif obsConf['type'] == 'Dummy':
|
|
||||||
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
|
||||||
return observables.Observable(**conf)
|
|
||||||
else:
|
|
||||||
raise Exception('Unknown Observable selected')
|
|
||||||
|
|
||||||
|
|
||||||
class ColumbusEnv(gym.Env):
|
class ColumbusEnv(gym.Env):
|
||||||
metadata = {'render.modes': ['human']}
|
metadata = {'render.modes': ['human'], 'render_modes': [
|
||||||
|
'human', 'non-human'], 'render_fps': 60}
|
||||||
|
|
||||||
def __init__(self, observable=observables.Observable(), fps=60, env_seed=3.1, master_seed=None, start_pos=(0.5, 0.5), start_score=0, speed_fac=0.01, acc_fac=0.04, die_on_zero=False, return_on_score=-1, reward_mult=1, agent_drag=0, controll_type='SPEED', aux_reward_max=1, aux_penalty_max=0, aux_reward_discretize=0, void_is_type_barrier=True, void_damage=1, torus_topology=False, default_collision_elasticity=1):
|
def __init__(self, observable=observables.Observable(), fps=60, env_seed=3.1, master_seed=None, start_pos=(0.5, 0.5), start_score=0, speed_fac=0.01, acc_fac=0.04, die_on_zero=False, return_on_score=-1, reward_mult=1, agent_drag=0, controll_type='SPEED', aux_reward_max=1, aux_penalty_max=0, aux_reward_discretize=0, void_is_type_barrier=True, void_damage=1, torus_topology=False, default_collision_elasticity=1, terminate_on_reward=False, agent_draw_path=False, clear_path_on_reset=True, max_steps=-1, value_color_mapper='tanh', width=720, height=720, agent_attrs={}, agent_cls=entities.Agent, exception_for_unsupported_collision=True, path_decay=0.1):
|
||||||
super(ColumbusEnv, self).__init__()
|
super(ColumbusEnv, self).__init__()
|
||||||
self.action_space = spaces.Box(
|
self.action_space = spaces.Box(
|
||||||
low=-1, high=1, shape=(2,), dtype=np.float32)
|
low=-1, high=1, shape=(2,), dtype=np.float32)
|
||||||
@@ -53,14 +24,14 @@ class ColumbusEnv(gym.Env):
|
|||||||
observable = parseObs(observable)
|
observable = parseObs(observable)
|
||||||
observable._set_env(self)
|
observable._set_env(self)
|
||||||
self.observable = observable
|
self.observable = observable
|
||||||
self.title = 'Untitled'
|
self.title = 'Columbus Env'
|
||||||
self.fps = fps
|
self.fps = fps
|
||||||
self.env_seed = env_seed
|
self.env_seed = env_seed
|
||||||
self.joystick_offset = (10, 10)
|
self.joystick_offset = (10, 10)
|
||||||
self.surface = None
|
self.surface = None
|
||||||
self.screen = None
|
self.screen = None
|
||||||
self.width = 720
|
self.width = width
|
||||||
self.height = 720
|
self.height = height
|
||||||
self.visible = False
|
self.visible = False
|
||||||
self.start_pos = start_pos
|
self.start_pos = start_pos
|
||||||
self.speed_fac = speed_fac/fps*60
|
self.speed_fac = speed_fac/fps*60
|
||||||
@@ -78,8 +49,7 @@ class ColumbusEnv(gym.Env):
|
|||||||
self.aux_penalty_max = aux_penalty_max # 0 = off
|
self.aux_penalty_max = aux_penalty_max # 0 = off
|
||||||
self.aux_reward_discretize = aux_reward_discretize
|
self.aux_reward_discretize = aux_reward_discretize
|
||||||
# 0 = dont discretize; how many steps (along diagonal)
|
# 0 = dont discretize; how many steps (along diagonal)
|
||||||
self.aux_reward_discretize = 0
|
self.penalty_from_edges = True # Don't change, only here to allow legacy behavior
|
||||||
self.penalty_from_edges = False
|
|
||||||
self.draw_observable = True
|
self.draw_observable = True
|
||||||
self.draw_joystick = True
|
self.draw_joystick = True
|
||||||
self.draw_entities = True
|
self.draw_entities = True
|
||||||
@@ -89,6 +59,27 @@ class ColumbusEnv(gym.Env):
|
|||||||
self.void_damage = void_damage
|
self.void_damage = void_damage
|
||||||
self.torus_topology = torus_topology
|
self.torus_topology = torus_topology
|
||||||
self.default_collision_elasticity = default_collision_elasticity
|
self.default_collision_elasticity = default_collision_elasticity
|
||||||
|
self.terminate_on_reward = terminate_on_reward
|
||||||
|
self.agent_draw_path = agent_draw_path
|
||||||
|
self.clear_path_on_reset = clear_path_on_reset
|
||||||
|
self.path_decay = path_decay
|
||||||
|
|
||||||
|
if isinstance(agent_cls, str):
|
||||||
|
agent_cls = getattr(entities, agent_cls)
|
||||||
|
self.Agent_cls = agent_cls
|
||||||
|
self.agent_attrs = agent_attrs
|
||||||
|
|
||||||
|
self.exception_for_unsupported_collision = exception_for_unsupported_collision
|
||||||
|
|
||||||
|
if value_color_mapper == 'atan':
|
||||||
|
def value_color_mapper(x): return th.atan(x*2)/0.786/2
|
||||||
|
elif value_color_mapper == 'tanh':
|
||||||
|
def value_color_mapper(x): return th.tanh(x*2)/0.762/2
|
||||||
|
self.value_color_mapper = value_color_mapper
|
||||||
|
|
||||||
|
self.max_steps = max_steps
|
||||||
|
self._steps = 0
|
||||||
|
self._has_value_map = False
|
||||||
|
|
||||||
self.paused = False
|
self.paused = False
|
||||||
self.keypress_timeout = 0
|
self.keypress_timeout = 0
|
||||||
@@ -104,6 +95,8 @@ class ColumbusEnv(gym.Env):
|
|||||||
|
|
||||||
self._init = False
|
self._init = False
|
||||||
|
|
||||||
|
self.is_columbus_env = True
|
||||||
|
|
||||||
@property
|
@property
|
||||||
def observation_space(self):
|
def observation_space(self):
|
||||||
if not self._init:
|
if not self._init:
|
||||||
@@ -121,6 +114,10 @@ class ColumbusEnv(gym.Env):
|
|||||||
def _ensure_surface(self):
|
def _ensure_surface(self):
|
||||||
if not self.surface or not self.screen:
|
if not self.surface or not self.screen:
|
||||||
self.surface = pygame.Surface((self.width, self.height))
|
self.surface = pygame.Surface((self.width, self.height))
|
||||||
|
self.path_overlay = pygame.Surface(
|
||||||
|
(self.width, self.height), pygame.SRCALPHA, 32)
|
||||||
|
self.value_overlay = pygame.Surface(
|
||||||
|
(self.width, self.height), pygame.SRCALPHA, 32)
|
||||||
if self.visible:
|
if self.visible:
|
||||||
self.screen = pygame.display.set_mode(
|
self.screen = pygame.display.set_mode(
|
||||||
(self.width, self.height))
|
(self.width, self.height))
|
||||||
@@ -171,9 +168,32 @@ class ColumbusEnv(gym.Env):
|
|||||||
elif isinstance(entity, entities.Enemy):
|
elif isinstance(entity, entities.Enemy):
|
||||||
if entity.radiateDamage:
|
if entity.radiateDamage:
|
||||||
if self.penalty_from_edges:
|
if self.penalty_from_edges:
|
||||||
penalty = self.aux_penalty_max / \
|
if self.agent.shape != 'circle':
|
||||||
(1 + self.sq_dist(entity.pos,
|
raise Exception(
|
||||||
self.agent.pos) - entity.radius - self.agent.redius)
|
'Radiating damage from edge for non-circle Agents not supported')
|
||||||
|
if entity.shape == 'circle':
|
||||||
|
penalty = self.aux_penalty_max / \
|
||||||
|
(1 + self.sq_dist(entity.pos,
|
||||||
|
self.agent.pos) - (entity.radius/max(self.height, self.width))**2 - (self.agent.radius/max(self.height, self.width))**2)
|
||||||
|
elif entity.shape == 'rect':
|
||||||
|
ax, ay = self.agent.pos
|
||||||
|
ex, ey, ex2, ey2 = entity.pos[0], entity.pos[1], entity.pos[0] + \
|
||||||
|
entity.width / \
|
||||||
|
self.width, entity.pos[1] + \
|
||||||
|
entity.height/self.height
|
||||||
|
lx, ly = ax, ay # 'Lotpunkt'
|
||||||
|
if ax < ex:
|
||||||
|
lx = ex
|
||||||
|
elif ax > ex2:
|
||||||
|
lx = ex2
|
||||||
|
if ay < ey:
|
||||||
|
ly = ey
|
||||||
|
elif ay > ey2:
|
||||||
|
ly = ey2
|
||||||
|
penalty = self.aux_penalty_max / \
|
||||||
|
(1 + self.sq_dist((lx, ly),
|
||||||
|
(ax, ay)) - (self.agent.radius/max(self.height, self.width))**2)
|
||||||
|
|
||||||
else:
|
else:
|
||||||
penalty = self.aux_penalty_max / \
|
penalty = self.aux_penalty_max / \
|
||||||
(1 + self.sq_dist(entity.pos, self.agent.pos))
|
(1 + self.sq_dist(entity.pos, self.agent.pos))
|
||||||
@@ -200,17 +220,20 @@ class ColumbusEnv(gym.Env):
|
|||||||
self._step_timers()
|
self._step_timers()
|
||||||
self._step_entities()
|
self._step_entities()
|
||||||
observation = self.observable.get_observation()
|
observation = self.observable.get_observation()
|
||||||
|
gotRew = self.new_reward > 0 or self.new_abs_reward > 0
|
||||||
|
self.gotHarm = self.new_reward < 0 or self.new_abs_reward < 0
|
||||||
reward, self.new_reward, self.new_abs_reward = self.new_reward / \
|
reward, self.new_reward, self.new_abs_reward = self.new_reward / \
|
||||||
self.fps + self.new_abs_reward, 0, 0
|
self.fps + self.new_abs_reward, 0, 0
|
||||||
if not self.torus_topology:
|
if not self.torus_topology:
|
||||||
if self.agent.pos[0] < 0.001 or self.agent.pos[0] > 0.999 \
|
if self.agent.getTop() < 1 or self.agent.getBottom() > self.height-1 \
|
||||||
or self.agent.pos[1] < 0.001 or self.agent.pos[1] > 0.999:
|
or self.agent.getLeft() < 1 or self.agent.getRight() > self.width-1:
|
||||||
reward -= self.void_damage/self.fps
|
reward -= self.void_damage/self.fps
|
||||||
self.score += reward # aux_reward does not count towards the score
|
self.score += reward # aux_reward does not count towards the score
|
||||||
if self.aux_reward_max:
|
if self.aux_reward_max or self.aux_penalty_max:
|
||||||
reward += self._get_aux_reward()
|
reward += self._get_aux_reward()
|
||||||
done = self.die_on_zero and self.score <= 0 or self.return_on_score != - \
|
self._steps += 1
|
||||||
1 and self.score > self.return_on_score
|
done = (self.die_on_zero and self.score <= 0) or (self.return_on_score != -
|
||||||
|
1 and self.score > self.return_on_score) or (self._steps == self.max_steps) or (self.terminate_on_reward and gotRew)
|
||||||
info = {'score': self.score, 'reward': reward}
|
info = {'score': self.score, 'reward': reward}
|
||||||
self._rendered = False
|
self._rendered = False
|
||||||
if done:
|
if done:
|
||||||
@@ -237,8 +260,10 @@ class ColumbusEnv(gym.Env):
|
|||||||
elif shapes == ['circle', 'rect']:
|
elif shapes == ['circle', 'rect']:
|
||||||
return sum([abs(d) for d in e1._get_crash_force_dir(e2)])
|
return sum([abs(d) for d in e1._get_crash_force_dir(e2)])
|
||||||
else:
|
else:
|
||||||
raise Exception(
|
if self.exception_for_unsupported_collision:
|
||||||
'Checking for collision between unsupported shapes: '+str(shapes))
|
raise Exception(
|
||||||
|
'Checking for collision between unsupported shapes: '+str(shapes))
|
||||||
|
return 0.0
|
||||||
|
|
||||||
def kill_entity(self, target):
|
def kill_entity(self, target):
|
||||||
newEntities = []
|
newEntities = []
|
||||||
@@ -254,9 +279,17 @@ class ColumbusEnv(gym.Env):
|
|||||||
self.agent.pos = self.start_pos
|
self.agent.pos = self.start_pos
|
||||||
# Expand this function
|
# Expand this function
|
||||||
|
|
||||||
def reset(self):
|
def _spawnAgent(self):
|
||||||
|
self.agent = self.Agent_cls(self)
|
||||||
|
self.agent.draw_path = self.agent_draw_path
|
||||||
|
for k, v in self.agent_attrs.items():
|
||||||
|
setattr(self.agent, k, v)
|
||||||
|
|
||||||
|
def reset(self, force_reset_path=False):
|
||||||
pygame.init()
|
pygame.init()
|
||||||
self._init = True
|
self._init = True
|
||||||
|
self._steps = 0
|
||||||
|
self._has_value_map = False
|
||||||
self._seed(self.env_seed)
|
self._seed(self.env_seed)
|
||||||
self._rendered = False
|
self._rendered = False
|
||||||
self._disturb_next = False
|
self._disturb_next = False
|
||||||
@@ -264,19 +297,68 @@ class ColumbusEnv(gym.Env):
|
|||||||
# will get rescaled acording to fps (=reward per second)
|
# will get rescaled acording to fps (=reward per second)
|
||||||
self.new_reward = 0
|
self.new_reward = 0
|
||||||
self.new_abs_reward = 0 # will not get rescaled. should be used for one-time rewards
|
self.new_abs_reward = 0 # will not get rescaled. should be used for one-time rewards
|
||||||
|
self.gotHarm = False
|
||||||
self.score = self.start_score
|
self.score = self.start_score
|
||||||
self.entities = []
|
self.entities = []
|
||||||
self.timers = []
|
self.timers = []
|
||||||
self.agent = entities.Agent(self)
|
self._spawnAgent()
|
||||||
self.setup()
|
self.setup()
|
||||||
self.entities.append(self.agent) # add it last, will be drawn on top
|
self.entities.append(self.agent) # add it last, will be drawn on top
|
||||||
self.observable.reset()
|
self.observable.reset()
|
||||||
|
if self.clear_path_on_reset or force_reset_path:
|
||||||
|
self._reset_paths()
|
||||||
return self.observable.get_observation()
|
return self.observable.get_observation()
|
||||||
|
|
||||||
|
def _reset_paths(self):
|
||||||
|
self.path_overlay = pygame.Surface(
|
||||||
|
(self.width, self.height), pygame.SRCALPHA, 32)
|
||||||
|
|
||||||
def _draw_entities(self):
|
def _draw_entities(self):
|
||||||
for entity in self.entities:
|
for entity in self.entities:
|
||||||
entity.draw()
|
entity.draw()
|
||||||
|
|
||||||
|
def _invalidate_value_map(self):
|
||||||
|
self._has_value_map = False
|
||||||
|
|
||||||
|
def _draw_values(self, value_func, static=True, resolution=64, color_depth=224, color_mapper=None):
|
||||||
|
if (not (static and self._has_value_map)):
|
||||||
|
agentpos = self.agent.pos
|
||||||
|
agentspeed = self.agent.speed
|
||||||
|
self.agent.speed = (0, 0)
|
||||||
|
self.value_overlay = pygame.Surface(
|
||||||
|
(self.width, self.height), pygame.SRCALPHA, 32)
|
||||||
|
obs = []
|
||||||
|
for i in range(resolution):
|
||||||
|
for j in range(resolution):
|
||||||
|
x, y = (i+0.5)/resolution, (j+0.5)/resolution
|
||||||
|
self.agent.pos = x, y
|
||||||
|
ob = self.observable.get_observation()
|
||||||
|
obs.append(ob)
|
||||||
|
self.agent.pos = agentpos
|
||||||
|
self.agent.speed = agentspeed
|
||||||
|
|
||||||
|
V = value_func(th.Tensor(np.array(obs)))
|
||||||
|
V /= max(V.max(), -1*V.min())*2
|
||||||
|
if color_mapper != None:
|
||||||
|
V = color_mapper(V)
|
||||||
|
V += 0.5
|
||||||
|
|
||||||
|
c = 0
|
||||||
|
for i in range(resolution):
|
||||||
|
for j in range(resolution):
|
||||||
|
v = V[c].item()
|
||||||
|
c += 1
|
||||||
|
col = [int((1-v)*color_depth),
|
||||||
|
int(v*color_depth), 0, color_depth]
|
||||||
|
x, y = i*(self.width/resolution), j * \
|
||||||
|
(self.height/resolution)
|
||||||
|
rect = pygame.Rect(x, y, int(self.width/resolution)+1,
|
||||||
|
int(self.height/resolution)+1)
|
||||||
|
pygame.draw.rect(self.value_overlay, col,
|
||||||
|
rect, width=0)
|
||||||
|
self.surface.blit(self.value_overlay, (0, 0))
|
||||||
|
self._has_value_map = True
|
||||||
|
|
||||||
def _draw_observable(self, forceDraw=False):
|
def _draw_observable(self, forceDraw=False):
|
||||||
if self.draw_observable and (self.visible or forceDraw):
|
if self.draw_observable and (self.visible or forceDraw):
|
||||||
self.observable.draw()
|
self.observable.draw()
|
||||||
@@ -293,13 +375,13 @@ class ColumbusEnv(gym.Env):
|
|||||||
pygame.draw.circle(self.screen, smolcol, (20+int(60*x) +
|
pygame.draw.circle(self.screen, smolcol, (20+int(60*x) +
|
||||||
self.joystick_offset[0], 20+int(60*y)+self.joystick_offset[1]), 20, width=0)
|
self.joystick_offset[0], 20+int(60*y)+self.joystick_offset[1]), 20, width=0)
|
||||||
|
|
||||||
def _draw_confidence_ellipse(self, chol, forceDraw=False, seconds=1):
|
def _draw_confidence_ellipse(self, chol, forceDraw=False, seconds=0.1):
|
||||||
# The 'seconds'-parameter only really makes sense, when using control_type='SPEED',
|
# The 'seconds'-parameter only really makes sense, when using control_type='SPEED',
|
||||||
# you can still use it to scale the cov-ellipse when using control_type='ACC',
|
# you can still use it to scale the cov-ellipse when using control_type='ACC',
|
||||||
# but it's relation to 'seconds' is no longer there...
|
# but it's relation to 'seconds' is no longer there...
|
||||||
if self.draw_confidence_ellipse and (self.visible or forceDraw):
|
if self.draw_confidence_ellipse and (self.visible or forceDraw):
|
||||||
col = (255, 255, 255)
|
col = (255, 255, 255)
|
||||||
f = seconds/self.speed_fac
|
f = seconds*self.speed_fac*self.fps*max(self.height, self.width)
|
||||||
|
|
||||||
while len(chol.shape) > 2:
|
while len(chol.shape) > 2:
|
||||||
chol = chol[0]
|
chol = chol[0]
|
||||||
@@ -311,21 +393,19 @@ class ColumbusEnv(gym.Env):
|
|||||||
|
|
||||||
L, V = th.linalg.eig(cov)
|
L, V = th.linalg.eig(cov)
|
||||||
L, V = L.real, V.real
|
L, V = L.real, V.real
|
||||||
w, h = int(abs(L[0].item()*f))+1, int(abs(L[1].item()*f))+1
|
l1, l2 = int(abs(math.sqrt(L[0].item())*f)) + \
|
||||||
# In theory we would have to solve:
|
1, int(abs(math.sqrt(L[1].item())*f))+1
|
||||||
# R = [[cos, -sin],[sin, cos]]
|
|
||||||
# But we only use the -sin term.
|
|
||||||
# Because of this our calculated angle might be wrong
|
|
||||||
# by periods of 180°
|
|
||||||
# But since an ellipsoid does not change under such an 'error',
|
|
||||||
# we don't care
|
|
||||||
# ang1 = int(math.acos(V[0, 0])/math.pi*360)
|
|
||||||
ang2 = int(math.asin(-V[0, 1])/math.pi*360)
|
|
||||||
# ang3 = int(math.asin(V[1, 0])/math.pi*360)
|
|
||||||
ang = ang2
|
|
||||||
|
|
||||||
# print(cov)
|
if l1 >= l2:
|
||||||
# print(w, h, (ang1, ang2, ang3))
|
w, h = l1, l2
|
||||||
|
run, rise = V[0][0], V[0][1]
|
||||||
|
else:
|
||||||
|
w, h = l2, l1
|
||||||
|
run, rise = V[1][0], V[1][1]
|
||||||
|
|
||||||
|
ang = (math.atan(rise/run))/(2*math.pi)*360
|
||||||
|
|
||||||
|
# print(w, h, (run, rise, ang))
|
||||||
|
|
||||||
x, y = self.agent.pos
|
x, y = self.agent.pos
|
||||||
x, y = x*self.width, y*self.height
|
x, y = x*self.width, y*self.height
|
||||||
@@ -337,6 +417,14 @@ class ColumbusEnv(gym.Env):
|
|||||||
self.screen.blit(rotated_surf, rotated_surf.get_rect(
|
self.screen.blit(rotated_surf, rotated_surf.get_rect(
|
||||||
center=rect.center))
|
center=rect.center))
|
||||||
|
|
||||||
|
def _draw_paths(self):
|
||||||
|
if self.path_decay != 0.0:
|
||||||
|
s = pygame.Surface((self.width, self.height))
|
||||||
|
s.set_alpha(soft_int(255*self.path_decay/self.fps))
|
||||||
|
s.fill((0, 0, 0))
|
||||||
|
self.path_overlay.blit(s, (0, 0))
|
||||||
|
self.surface.blit(self.path_overlay, (0, 0))
|
||||||
|
|
||||||
def _handle_user_input(self):
|
def _handle_user_input(self):
|
||||||
for event in pygame.event.get():
|
for event in pygame.event.get():
|
||||||
pass
|
pass
|
||||||
@@ -349,6 +437,8 @@ class ColumbusEnv(gym.Env):
|
|||||||
self.draw_confidence_ellipse = not self.draw_confidence_ellipse
|
self.draw_confidence_ellipse = not self.draw_confidence_ellipse
|
||||||
elif keys[pygame.K_r]:
|
elif keys[pygame.K_r]:
|
||||||
self.reset()
|
self.reset()
|
||||||
|
elif keys[pygame.K_t]:
|
||||||
|
self._reset_paths()
|
||||||
elif keys[pygame.K_p]:
|
elif keys[pygame.K_p]:
|
||||||
self.paused = not self.paused
|
self.paused = not self.paused
|
||||||
else:
|
else:
|
||||||
@@ -369,13 +459,17 @@ class ColumbusEnv(gym.Env):
|
|||||||
elif keys[pygame.K_d]:
|
elif keys[pygame.K_d]:
|
||||||
self._disturb_next = (1.0, 0.5)
|
self._disturb_next = (1.0, 0.5)
|
||||||
|
|
||||||
def render(self, mode='human', dont_show=False, chol=None):
|
def render(self, mode='human', dont_show=False, chol=None, value_func=None, values_static=True):
|
||||||
if mode == 'human':
|
if mode == 'human':
|
||||||
self._handle_user_input()
|
self._handle_user_input()
|
||||||
self.visible = self.visible or not dont_show
|
self.visible = self.visible or not dont_show
|
||||||
self._ensure_surface()
|
self._ensure_surface()
|
||||||
pygame.draw.rect(self.surface, (0, 0, 0),
|
pygame.draw.rect(self.surface, (0, 0, 0),
|
||||||
pygame.Rect(0, 0, self.width, self.height))
|
pygame.Rect(0, 0, self.width, self.height))
|
||||||
|
if value_func != None:
|
||||||
|
self._draw_values(value_func, values_static,
|
||||||
|
color_mapper=self.value_color_mapper)
|
||||||
|
self._draw_paths()
|
||||||
if self.draw_entities:
|
if self.draw_entities:
|
||||||
self._draw_entities()
|
self._draw_entities()
|
||||||
else:
|
else:
|
||||||
@@ -398,6 +492,123 @@ class ColumbusEnv(gym.Env):
|
|||||||
pygame.quit()
|
pygame.quit()
|
||||||
|
|
||||||
|
|
||||||
|
class ColumbusConfigDefined(ColumbusEnv):
|
||||||
|
# Allows defining Columbus Environments using dicts.
|
||||||
|
# Intended to be used in combination with cw2 configuration.
|
||||||
|
# Look into humanPlayer to see how this is supposed to be interfaced with.
|
||||||
|
|
||||||
|
def __init__(self, observable={}, env_seed=None, entities=[], fps=30, **kw):
|
||||||
|
super().__init__(
|
||||||
|
observable=observable, fps=fps, env_seed=env_seed, **kw)
|
||||||
|
self.entities_definitions = entities
|
||||||
|
self.start_pos = self.conv_unit(self.start_pos[0], target='em', axis='x'), self.conv_unit(
|
||||||
|
self.start_pos[1], target='em', axis='y')
|
||||||
|
|
||||||
|
def is_unit(self, s):
|
||||||
|
if type(s) in [int, float]:
|
||||||
|
return True
|
||||||
|
if s.replace('.', '', 1).replace('-', '0', 1).isdigit():
|
||||||
|
return True
|
||||||
|
num, unit = s[:-2], s[-2:]
|
||||||
|
if unit in ['px', 'em', 'rx', 'ry', 'ct', 'au']:
|
||||||
|
if num.replace('.', '', 1).replace('-', '0', 1).isdigit():
|
||||||
|
return True
|
||||||
|
return False
|
||||||
|
|
||||||
|
def conv_unit(self, s, target='px', axis='x'):
|
||||||
|
assert self.is_unit(s)
|
||||||
|
if type(s) in [int, float]:
|
||||||
|
return s
|
||||||
|
if s.replace('.', '', 1).isdigit():
|
||||||
|
if target == 'px':
|
||||||
|
return int(s)
|
||||||
|
return float(s)
|
||||||
|
num, unit = s[:-2], s[-2:]
|
||||||
|
num = float(num)
|
||||||
|
if unit == 'rx':
|
||||||
|
unit = 'px'
|
||||||
|
axis = 'x'
|
||||||
|
elif unit == 'ry':
|
||||||
|
unit = 'px'
|
||||||
|
axis = 'y'
|
||||||
|
if unit == 'em':
|
||||||
|
em = num
|
||||||
|
elif unit == 'px':
|
||||||
|
em = num / ({'x': self.width, 'y': self.height}[axis])
|
||||||
|
elif unit == 'au':
|
||||||
|
em = num * 36 / ({'x': self.width, 'y': self.height}[axis])
|
||||||
|
elif unit == 'ct':
|
||||||
|
em = num / 100
|
||||||
|
else:
|
||||||
|
raise Exception('Conversion not implemented')
|
||||||
|
|
||||||
|
if target == 'em':
|
||||||
|
return em
|
||||||
|
elif target == 'px':
|
||||||
|
return int(em * ({'x': self.width, 'y': self.height}[axis]))
|
||||||
|
|
||||||
|
def setup(self):
|
||||||
|
self.agent.pos = self.start_pos
|
||||||
|
for i, e in enumerate(self.entities_definitions):
|
||||||
|
Entity = getattr(entities, e['type'])
|
||||||
|
for i in range(e.get('num', 1) + int(self.random()*(0.99+e.get('num_rand', 0)))):
|
||||||
|
entity = Entity(self)
|
||||||
|
conf = {k: v for k, v in e.items() if str(
|
||||||
|
k) not in ['num', 'num_rand', 'type']}
|
||||||
|
|
||||||
|
for k, v_raw in conf.items():
|
||||||
|
if k == 'pos':
|
||||||
|
v = self.conv_unit(v_raw[0], target='em', axis='x'), self.conv_unit(
|
||||||
|
v_raw[1], target='em', axis='y')
|
||||||
|
elif k in ['width', 'height', 'radius']:
|
||||||
|
v = self.conv_unit(
|
||||||
|
v_raw, target='px', axis='y' if k == 'height' else 'x')
|
||||||
|
else:
|
||||||
|
v = v_raw
|
||||||
|
if k.endswith('_rand'):
|
||||||
|
v = self.conv_unit(
|
||||||
|
v_raw, target='px', axis='y' if k == 'height_rand' else 'x')
|
||||||
|
if isinstance(v, int):
|
||||||
|
n = k.replace('_rand', '')
|
||||||
|
cur = getattr(
|
||||||
|
entity, n)
|
||||||
|
inc = int((v+0.99)*self.random())
|
||||||
|
setattr(entity, n, cur + inc)
|
||||||
|
elif isinstance(v, float):
|
||||||
|
n = k.replace('_rand', '')
|
||||||
|
cur = getattr(
|
||||||
|
entity, n)
|
||||||
|
inc = v*self.random()
|
||||||
|
setattr(entity, n, cur + inc)
|
||||||
|
elif isinstance(v, list):
|
||||||
|
for vi, ve in enumerate(v):
|
||||||
|
if isinstance(v, int):
|
||||||
|
n = k.replace('_rand', '')
|
||||||
|
cur = getattr(
|
||||||
|
entity, n)
|
||||||
|
cur[vi] = int((v+0.99)*self.random())
|
||||||
|
setattr(entity, n, cur)
|
||||||
|
elif isinstance(v, float):
|
||||||
|
n = k.replace('_rand', '')
|
||||||
|
cur = getattr(
|
||||||
|
entity, n)
|
||||||
|
cur[vi] = v*self.random()
|
||||||
|
setattr(entity, n, cur)
|
||||||
|
elif k.endswith('_randf'):
|
||||||
|
n = k.replace('_randf', '')
|
||||||
|
cur = getattr(
|
||||||
|
entity, n)
|
||||||
|
inc = v*self.random()
|
||||||
|
setattr(entity, n, cur + inc)
|
||||||
|
else:
|
||||||
|
setattr(entity, k, v)
|
||||||
|
|
||||||
|
self.entities.append(entity)
|
||||||
|
|
||||||
|
###
|
||||||
|
# Custom Env Definitions
|
||||||
|
|
||||||
|
|
||||||
class ColumbusTest3_1(ColumbusEnv):
|
class ColumbusTest3_1(ColumbusEnv):
|
||||||
def __init__(self, observable=observables.CnnObservable(out_width=48, out_height=48), fps=30, aux_reward_max=1, **kw):
|
def __init__(self, observable=observables.CnnObservable(out_width=48, out_height=48), fps=30, aux_reward_max=1, **kw):
|
||||||
super(ColumbusTest3_1, self).__init__(
|
super(ColumbusTest3_1, self).__init__(
|
||||||
@@ -737,59 +948,35 @@ class ColumbusFootball(ColumbusEnv):
|
|||||||
self.entities.append(entities.FlyingFootballPlayer(self, ball))
|
self.entities.append(entities.FlyingFootballPlayer(self, ball))
|
||||||
|
|
||||||
|
|
||||||
class ColumbusConfigDefined(ColumbusEnv):
|
|
||||||
def __init__(self, observable={}, env_seed=None, entities=[], fps=30, **kw):
|
|
||||||
super().__init__(
|
|
||||||
observable=observable, fps=fps, env_seed=env_seed, **kw)
|
|
||||||
self.entities_definitions = entities
|
|
||||||
|
|
||||||
def setup(self):
|
|
||||||
self.agent.pos = self.start_pos
|
|
||||||
for i, e in enumerate(self.entities_definitions):
|
|
||||||
Entity = getattr(entities, e['type'])
|
|
||||||
for i in range(e.get('num', 1) + int(self.random()*(0.99+e.get('num_rand', 0)))):
|
|
||||||
entity = Entity(self)
|
|
||||||
conf = {k: v for k, v in e.items() if str(
|
|
||||||
k) not in ['num', 'num_rand', 'type']}
|
|
||||||
|
|
||||||
for k, v in conf.items():
|
|
||||||
if k.endswith('_rand'):
|
|
||||||
n = k.replace('_rand', '')
|
|
||||||
cur = getattr(
|
|
||||||
entity, n)
|
|
||||||
inc = int((v+0.99)*self.random())
|
|
||||||
setattr(entity, n, cur + inc)
|
|
||||||
elif k.endswith('_randf'):
|
|
||||||
n = k.replace('_randf', '')
|
|
||||||
cur = getattr(
|
|
||||||
entity, n)
|
|
||||||
inc = v*self.random()
|
|
||||||
setattr(entity, n, cur + inc)
|
|
||||||
else:
|
|
||||||
setattr(entity, k, v)
|
|
||||||
|
|
||||||
self.entities.append(entity)
|
|
||||||
|
|
||||||
|
|
||||||
class ColumbusBlub(ColumbusEnv):
|
class ColumbusBlub(ColumbusEnv):
|
||||||
def __init__(self, observable=observables.CompositionalObservable([observables.StateObservable(), observables.RayObservable(num_rays=6, chans=[entities.Enemy])]), env_seed=None, entities=[], fps=30, **kw):
|
def __init__(self, observable=observables.CompositionalObservable([observables.StateObservable(), observables.RayObservable(num_rays=6, chans=[entities.Enemy])]), env_seed=None, entities=[], fps=30, **kw):
|
||||||
super().__init__(
|
super().__init__(
|
||||||
observable=observable, fps=fps, env_seed=env_seed, default_collision_elasticity=0.8, speed_fac=0.01, acc_fac=0.1, agent_drag=0.06, controll_type='ACC')
|
observable=observable, fps=fps, env_seed=env_seed, default_collision_elasticity=0.8, speed_fac=0.01, acc_fac=0.1, agent_drag=0.06, controll_type='ACC', aux_penalty_max=1)
|
||||||
|
|
||||||
def setup(self):
|
def setup(self):
|
||||||
self.agent.pos = self.start_pos
|
self.agent.pos = self.start_pos
|
||||||
for i in range(10):
|
|
||||||
enemy = entities.CircleBarrier(self)
|
|
||||||
enemy.radius = self.random()*25+75
|
|
||||||
self.entities.append(enemy)
|
|
||||||
for i in range(1):
|
for i in range(1):
|
||||||
reward = entities.TeleportingReward(self)
|
enemy = entities.RectBarrier(self)
|
||||||
reward.radius = 20
|
enemy.radius = 100
|
||||||
reward.reward = 25
|
enemy.width, enemy.height = 200, 75
|
||||||
self.entities.append(reward)
|
self.entities.append(enemy)
|
||||||
|
|
||||||
|
|
||||||
###
|
###
|
||||||
|
# Registering Envs fro Gym
|
||||||
|
register( # Legacy
|
||||||
|
id='ColumbusConfigDefined-v0',
|
||||||
|
entry_point=ColumbusConfigDefined,
|
||||||
|
max_episode_steps=30*60*2, # 2 min at default (30) fps
|
||||||
|
)
|
||||||
|
|
||||||
|
register(
|
||||||
|
id='Columbus-v1',
|
||||||
|
entry_point=ColumbusConfigDefined
|
||||||
|
)
|
||||||
|
|
||||||
|
###
|
||||||
|
|
||||||
# register(
|
# register(
|
||||||
# id='ColumbusBlub-v0',
|
# id='ColumbusBlub-v0',
|
||||||
# entry_point=ColumbusBlub,
|
# entry_point=ColumbusBlub,
|
||||||
@@ -797,17 +984,17 @@ class ColumbusBlub(ColumbusEnv):
|
|||||||
# )
|
# )
|
||||||
|
|
||||||
|
|
||||||
register(
|
# register(
|
||||||
id='ColumbusTestCnn-v0',
|
# id='ColumbusTestCnn-v0',
|
||||||
entry_point=ColumbusTest3_1,
|
# entry_point=ColumbusTest3_1,
|
||||||
max_episode_steps=30*60*2,
|
# max_episode_steps=30*60*2,
|
||||||
)
|
# )
|
||||||
|
|
||||||
register(
|
# register(
|
||||||
id='ColumbusTestRay-v0',
|
# id='ColumbusTestRay-v0',
|
||||||
entry_point=ColumbusTestRay,
|
# entry_point=ColumbusTestRay,
|
||||||
max_episode_steps=30*60*2,
|
# max_episode_steps=30*60*2,
|
||||||
)
|
# )
|
||||||
|
|
||||||
# register(
|
# register(
|
||||||
# id='ColumbusRayDrone-v0',
|
# id='ColumbusRayDrone-v0',
|
||||||
@@ -845,11 +1032,11 @@ register(
|
|||||||
# max_episode_steps=30*60*2,
|
# max_episode_steps=30*60*2,
|
||||||
# )
|
# )
|
||||||
|
|
||||||
register(
|
# register(
|
||||||
id='ColumbusStateWithBarriers-v0',
|
# id='ColumbusStateWithBarriers-v0',
|
||||||
entry_point=ColumbusStateWithBarriers,
|
# entry_point=ColumbusStateWithBarriers,
|
||||||
max_episode_steps=30*60*2,
|
# max_episode_steps=30*60*2,
|
||||||
)
|
# )
|
||||||
|
|
||||||
# register(
|
# register(
|
||||||
# id='ColumbusCompassWithBarriers-v0',
|
# id='ColumbusCompassWithBarriers-v0',
|
||||||
@@ -881,12 +1068,6 @@ register(
|
|||||||
# max_episode_steps=30*60*2,
|
# max_episode_steps=30*60*2,
|
||||||
# )
|
# )
|
||||||
|
|
||||||
register(
|
|
||||||
id='ColumbusConfigDefined-v0',
|
|
||||||
entry_point=ColumbusConfigDefined,
|
|
||||||
max_episode_steps=30*60*2,
|
|
||||||
)
|
|
||||||
|
|
||||||
register(
|
register(
|
||||||
id='ColumbusDemoEnvFootball-v0',
|
id='ColumbusDemoEnvFootball-v0',
|
||||||
entry_point=ColumbusDemoEnvFootball,
|
entry_point=ColumbusDemoEnvFootball,
|
||||||
|
|||||||
+43
-5
@@ -1,16 +1,18 @@
|
|||||||
|
import torch as th
|
||||||
from time import sleep, time
|
from time import sleep, time
|
||||||
import numpy as np
|
import numpy as np
|
||||||
import pygame
|
import pygame
|
||||||
|
import yaml
|
||||||
|
|
||||||
from columbus import env
|
from columbus import env
|
||||||
from columbus.observables import Observable, CnnObservable
|
from columbus.observables import Observable, CnnObservable
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
Env = chooseEnv()
|
env = chooseEnv()
|
||||||
env = Env(fps=30)
|
while True:
|
||||||
env.start_pos = [0.6, 0.3]
|
playEnv(env)
|
||||||
playEnv(env)
|
input('<again?>')
|
||||||
env.close()
|
env.close()
|
||||||
|
|
||||||
|
|
||||||
@@ -22,6 +24,33 @@ def getAvaibleEnvs():
|
|||||||
yield getattr(env, s)
|
yield getattr(env, s)
|
||||||
|
|
||||||
|
|
||||||
|
def loadConfigDefinedEnv(EnvClass):
|
||||||
|
p = input('[Path to config> ')
|
||||||
|
with open(p, 'r') as f:
|
||||||
|
docs = list([d for d in yaml.safe_load_all(
|
||||||
|
f) if d and 'name' in d and d['name'] not in ['SLURM']])
|
||||||
|
for i, doc in enumerate(docs):
|
||||||
|
name = doc['name']
|
||||||
|
print('['+str(i)+'] '+name)
|
||||||
|
ds = int(input('[0]> ') or '0')
|
||||||
|
doc = docs[ds]
|
||||||
|
cur = doc
|
||||||
|
path = 'params.task.env_args'
|
||||||
|
p = path.split('.')
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
if len(p) == 0:
|
||||||
|
break
|
||||||
|
key = p.pop(0)
|
||||||
|
print(key)
|
||||||
|
cur = cur[key]
|
||||||
|
except Exception as e:
|
||||||
|
print('Unable to find key "'+key+'"')
|
||||||
|
path = input('[Path> ')
|
||||||
|
print(cur)
|
||||||
|
return EnvClass(fps=30, **cur)
|
||||||
|
|
||||||
|
|
||||||
def chooseEnv():
|
def chooseEnv():
|
||||||
envs = list(getAvaibleEnvs())
|
envs = list(getAvaibleEnvs())
|
||||||
for i, Env in enumerate(envs):
|
for i, Env in enumerate(envs):
|
||||||
@@ -35,7 +64,15 @@ def chooseEnv():
|
|||||||
if i < 0 or i >= len(envs):
|
if i < 0 or i >= len(envs):
|
||||||
print(
|
print(
|
||||||
'[!] That is a number, but not one that makes sense in this context...')
|
'[!] That is a number, but not one that makes sense in this context...')
|
||||||
return envs[i]
|
if envs[i] in [env.ColumbusConfigDefined]:
|
||||||
|
return loadConfigDefinedEnv(envs[i])
|
||||||
|
Env = envs[i]
|
||||||
|
return Env(fps=30)
|
||||||
|
|
||||||
|
|
||||||
|
def value_func(obs):
|
||||||
|
return obs[:, 0]
|
||||||
|
# return th.rand(obs.shape[0])-0.5
|
||||||
|
|
||||||
|
|
||||||
def playEnv(env):
|
def playEnv(env):
|
||||||
@@ -43,6 +80,7 @@ def playEnv(env):
|
|||||||
env.reset()
|
env.reset()
|
||||||
while not done:
|
while not done:
|
||||||
t1 = time()
|
t1 = time()
|
||||||
|
# env.render(value_func=value_func)
|
||||||
env.render()
|
env.render()
|
||||||
pos = (0.5, 0.5)
|
pos = (0.5, 0.5)
|
||||||
pos = pygame.mouse.get_pos()
|
pos = pygame.mouse.get_pos()
|
||||||
|
|||||||
+16
-9
@@ -16,7 +16,7 @@ class Observable():
|
|||||||
def get_observation_space(self):
|
def get_observation_space(self):
|
||||||
print("[!] Using dummyObservable. Env won't output anything")
|
print("[!] Using dummyObservable. Env won't output anything")
|
||||||
return spaces.Box(low=0, high=1,
|
return spaces.Box(low=0, high=1,
|
||||||
shape=(1,), dtype=np.float32)
|
shape=(1,), dtype=np.float64)
|
||||||
|
|
||||||
def get_observation(self):
|
def get_observation(self):
|
||||||
return np.array([0])
|
return np.array([0])
|
||||||
@@ -45,7 +45,7 @@ class CnnObservable(Observable):
|
|||||||
|
|
||||||
def get_observation_space(self):
|
def get_observation_space(self):
|
||||||
return spaces.Box(low=0, high=255,
|
return spaces.Box(low=0, high=255,
|
||||||
shape=(self.out_width, self.out_height, 3), dtype=np.float32)
|
shape=(self.out_width, self.out_height, 3), dtype=np.float64)
|
||||||
|
|
||||||
def get_observation(self):
|
def get_observation(self):
|
||||||
if not self.env._rendered:
|
if not self.env._rendered:
|
||||||
@@ -132,24 +132,31 @@ class RayObservable(Observable):
|
|||||||
'Can only raycast circular and rectangular entities!')
|
'Can only raycast circular and rectangular entities!')
|
||||||
return False
|
return False
|
||||||
|
|
||||||
|
# Filter out entities, that we sure are out of range
|
||||||
|
# (so we have to do less work for the ray collisions)
|
||||||
def _get_possible_entities(self):
|
def _get_possible_entities(self):
|
||||||
entities_l = []
|
entities_l = []
|
||||||
if entities.Void in self.chans or self.env.void_barrier:
|
if entities.Void in self.chans or self.env.void_barrier:
|
||||||
entities_l.append(entities.Void(self.env))
|
entities_l.append(entities.Void(self.env))
|
||||||
for entity in self.env.entities:
|
for entity in self.env.entities:
|
||||||
if entity.shape == 'rect':
|
if entity.shape == 'rect':
|
||||||
|
x, y = entity.pos[0]+entity.width/self.env.width / \
|
||||||
|
2, entity.pos[1]+entity.height/self.env.height/2
|
||||||
radius = (entity.width/2 + entity.height/2)*1.0
|
radius = (entity.width/2 + entity.height/2)*1.0
|
||||||
elif entity.shape == 'circle':
|
elif entity.shape == 'circle':
|
||||||
|
x, y = entity.pos[0], entity.pos[1]
|
||||||
radius = entity.radius
|
radius = entity.radius
|
||||||
else:
|
else:
|
||||||
raise Exception(
|
raise Exception(
|
||||||
'Can only raycast circular and rectangular entities!')
|
'Can only raycast circular and rectangular entities!')
|
||||||
sq_dist = ((self.env.agent.pos[0]-entity.pos[0])*self.env.width) ** 2 \
|
sq_dist = ((self.env.agent.pos[0]-x)*self.env.width) ** 2 \
|
||||||
+ ((self.env.agent.pos[1]-entity.pos[1])*self.env.height) ** 2
|
+ ((self.env.agent.pos[1]-y)*self.env.height) ** 2
|
||||||
if sq_dist <= (radius + self.env.agent.radius + self.ray_len)**2:
|
if sq_dist <= (radius + self.env.agent.getQuasiRadius() + self.ray_len)**2:
|
||||||
entities_l.append(entity) # cannot use yield here!
|
entities_l.append(entity) # cannot use yield here!
|
||||||
return entities_l
|
return entities_l
|
||||||
|
|
||||||
|
# Ugly, inefficient ray casting
|
||||||
|
# Oh well, it works...
|
||||||
def get_observation(self):
|
def get_observation(self):
|
||||||
entities = self._get_possible_entities()
|
entities = self._get_possible_entities()
|
||||||
self.rays = np.zeros((self.num_rays+self.include_rand, self.num_chans))
|
self.rays = np.zeros((self.num_rays+self.include_rand, self.num_chans))
|
||||||
@@ -243,8 +250,8 @@ class StateObservable(Observable):
|
|||||||
self.reset()
|
self.reset()
|
||||||
num = len(self.entities)*2+len(self._timeoutEntities) + \
|
num = len(self.entities)*2+len(self._timeoutEntities) + \
|
||||||
self.speedAgent*2 + self.include_rand
|
self.speedAgent*2 + self.include_rand
|
||||||
return spaces.Box(low=0-1*self.coordsRelativeToAgent, high=1,
|
return spaces.Box(low=0-1*(self.coordsRelativeToAgent or self.speedAgent), high=1,
|
||||||
shape=(num,), dtype=np.float32)
|
shape=(num,), dtype=np.float64)
|
||||||
|
|
||||||
def get_observation(self):
|
def get_observation(self):
|
||||||
obs = []
|
obs = []
|
||||||
@@ -324,7 +331,7 @@ class CompassObservable(Observable):
|
|||||||
self.reset()
|
self.reset()
|
||||||
num = len(self.entities)*2
|
num = len(self.entities)*2
|
||||||
return spaces.Box(low=-1, high=1,
|
return spaces.Box(low=-1, high=1,
|
||||||
shape=(num,), dtype=np.float32)
|
shape=(num,), dtype=np.float64)
|
||||||
|
|
||||||
def reset(self):
|
def reset(self):
|
||||||
self._entities = None
|
self._entities = None
|
||||||
@@ -378,7 +385,7 @@ class CompositionalObservable(Observable):
|
|||||||
low = np.hstack((low, space.low.reshape((-1))))
|
low = np.hstack((low, space.low.reshape((-1))))
|
||||||
high = np.hstack((high, space.high.reshape((-1))))
|
high = np.hstack((high, space.high.reshape((-1))))
|
||||||
return spaces.Box(low=low, high=high,
|
return spaces.Box(low=low, high=high,
|
||||||
shape=(num,), dtype=np.float32)
|
shape=(num,), dtype=np.float64)
|
||||||
|
|
||||||
def get_observation(self):
|
def get_observation(self):
|
||||||
o = [obs.get_observation().reshape((-1))
|
o = [obs.get_observation().reshape((-1))
|
||||||
|
|||||||
@@ -0,0 +1,42 @@
|
|||||||
|
from columbus import entities, observables
|
||||||
|
|
||||||
|
import random as random_dont_use
|
||||||
|
|
||||||
|
|
||||||
|
def parseObs(obsConf):
|
||||||
|
# Parsing Observable Definitions
|
||||||
|
if type(obsConf) == list:
|
||||||
|
obs = []
|
||||||
|
for i, c in enumerate(obsConf):
|
||||||
|
obs.append(parseObs(c))
|
||||||
|
if len(obs) == 1:
|
||||||
|
return obs[0]
|
||||||
|
else:
|
||||||
|
return observables.CompositionalObservable(obs)
|
||||||
|
|
||||||
|
if obsConf['type'] == 'State':
|
||||||
|
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
||||||
|
return observables.StateObservable(**conf)
|
||||||
|
elif obsConf['type'] == 'Compass':
|
||||||
|
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
||||||
|
return observables.CompassObservable(**conf)
|
||||||
|
elif obsConf['type'] == 'RayCast':
|
||||||
|
chans = []
|
||||||
|
for chan in obsConf.get('chans', []):
|
||||||
|
chans.append(getattr(entities, chan))
|
||||||
|
conf = {k: v for k, v in obsConf.items() if k not in ['type', 'chans']}
|
||||||
|
return observables.RayObservable(chans=chans, **conf)
|
||||||
|
elif obsConf['type'] == 'CNN':
|
||||||
|
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
||||||
|
return observables.CnnObservable(**conf)
|
||||||
|
elif obsConf['type'] == 'Dummy':
|
||||||
|
conf = {k: v for k, v in obsConf.items() if k not in ['type']}
|
||||||
|
return observables.Observable(**conf)
|
||||||
|
else:
|
||||||
|
raise Exception('Unknown Observable selected')
|
||||||
|
|
||||||
|
|
||||||
|
def soft_int(num):
|
||||||
|
i = int(num)
|
||||||
|
r = num - i
|
||||||
|
return i + int(random_dont_use.random() < r)
|
||||||
@@ -0,0 +1,78 @@
|
|||||||
|
name: "DEFAULT"
|
||||||
|
|
||||||
|
params:
|
||||||
|
task:
|
||||||
|
task: columbus
|
||||||
|
env_name: ColumbusConfigDefined-v0
|
||||||
|
env_args:
|
||||||
|
observable:
|
||||||
|
- type: State
|
||||||
|
coordsAgent: True
|
||||||
|
speedAgent: True
|
||||||
|
coordsRelativeToAgent: False
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: State
|
||||||
|
coordsAgent: False
|
||||||
|
speedAgent: False
|
||||||
|
coordsRelativeToAgent: True
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: Compass
|
||||||
|
- type: RayCast
|
||||||
|
num_rays: 6
|
||||||
|
chans: [Enemy]
|
||||||
|
entities:
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 1 #1
|
||||||
|
width: 300
|
||||||
|
height: 120 # 360 - 5%(720)
|
||||||
|
pos: [0, 0]
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 1 #1
|
||||||
|
width: 300
|
||||||
|
height: 1000
|
||||||
|
pos: [0, 0.25]
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 1 #1
|
||||||
|
width: 250
|
||||||
|
height: 30
|
||||||
|
pos: [0.55, 0.6]
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 1 #1
|
||||||
|
width: 30
|
||||||
|
height: 120
|
||||||
|
pos: [0.856, 0.475]
|
||||||
|
- type: RectBarrier
|
||||||
|
num: 0
|
||||||
|
damage: 1 #1
|
||||||
|
width: 50
|
||||||
|
width_rand: 100
|
||||||
|
height: 25
|
||||||
|
height_rand: 100
|
||||||
|
- type: OnceReward
|
||||||
|
reward: 100
|
||||||
|
radius: 20
|
||||||
|
pos: [0.9, 0.8]
|
||||||
|
start_pos: [0.1, 0.21]
|
||||||
|
default_collision_elasticity: 0.8
|
||||||
|
start_score: 10
|
||||||
|
speed_fac: 0.01
|
||||||
|
acc_fac: 0.1
|
||||||
|
die_on_zero: False #True
|
||||||
|
agent_drag: 0.1 # 0.05
|
||||||
|
controll_type: ACC # SPEED
|
||||||
|
aux_reward_max: 1
|
||||||
|
aux_penalty_max: 0.01
|
||||||
|
void_damage: 5 #1
|
||||||
|
terminate_on_reward: True
|
||||||
|
agent_draw_path: True
|
||||||
|
clear_path_on_reset: False
|
||||||
|
max_steps: 450 # 1800
|
||||||
|
---
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
name: "DEFAULT"
|
||||||
|
|
||||||
|
params:
|
||||||
|
task:
|
||||||
|
task: columbus
|
||||||
|
num_envs: 8
|
||||||
|
env_args:
|
||||||
|
observable:
|
||||||
|
- type: State
|
||||||
|
coordsAgent: True
|
||||||
|
speedAgent: True
|
||||||
|
coordsRelativeToAgent: False
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: State
|
||||||
|
coordsAgent: False
|
||||||
|
speedAgent: False
|
||||||
|
coordsRelativeToAgent: True
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: Compass
|
||||||
|
- type: RayCast
|
||||||
|
num_rays: 8
|
||||||
|
chans: [Enemy]
|
||||||
|
entities:
|
||||||
|
- type: CircleBarrier
|
||||||
|
num: 8
|
||||||
|
num_rand: 6
|
||||||
|
damage: 20 #20
|
||||||
|
radius: 25
|
||||||
|
radius_rand: 75
|
||||||
|
- type: TeleportingReward
|
||||||
|
num: 1
|
||||||
|
reward: 100 #100
|
||||||
|
radius: 20
|
||||||
|
default_collision_elasticity: 0.8
|
||||||
|
start_score: 50
|
||||||
|
speed_fac: 0.01
|
||||||
|
acc_fac: 0.1
|
||||||
|
die_on_zero: True
|
||||||
|
agent_drag: 0.07 # 0.05
|
||||||
|
controll_type: ACC # SPEED
|
||||||
|
aux_reward_max: 1
|
||||||
|
aux_penalty_max: 0.1
|
||||||
|
void_damage: 5 #1
|
||||||
|
#master_seed: 3.14
|
||||||
|
max_steps: 900 # 30 sec
|
||||||
|
---
|
||||||
@@ -0,0 +1,67 @@
|
|||||||
|
name: "DEFAULT"
|
||||||
|
|
||||||
|
params:
|
||||||
|
task:
|
||||||
|
task: columbus
|
||||||
|
env_name: ColumbusConfigDefined-v0
|
||||||
|
env_args:
|
||||||
|
observable:
|
||||||
|
- type: State
|
||||||
|
coordsAgent: True
|
||||||
|
speedAgent: True
|
||||||
|
coordsRelativeToAgent: False
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: State
|
||||||
|
coordsAgent: False
|
||||||
|
speedAgent: False
|
||||||
|
coordsRelativeToAgent: True
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: RayCast
|
||||||
|
num_rays: 6
|
||||||
|
chans: [Enemy]
|
||||||
|
entities:
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 10 #1
|
||||||
|
width: 25
|
||||||
|
height: 120 # 360 - 5%(720)
|
||||||
|
pos: [0.45, 0]
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 10 #1
|
||||||
|
width: 25
|
||||||
|
height: 1000
|
||||||
|
pos: [0.45, 0.25]
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 10 #1
|
||||||
|
width: 25
|
||||||
|
height: 520 # 360 - 5%(720)
|
||||||
|
pos: [0.55, 0]
|
||||||
|
- type: RectBarrier
|
||||||
|
damage: 10 #1
|
||||||
|
width: 25
|
||||||
|
height: 200
|
||||||
|
pos: [0.55, 0.80]
|
||||||
|
- type: LoopReward
|
||||||
|
num: 1
|
||||||
|
reward: 100 #25
|
||||||
|
radius: 20
|
||||||
|
loop: [[0.125, 0.5, 0.1, 0.5], [0.875, 0.5, 0.1, 0.5]]
|
||||||
|
default_collision_elasticity: 0.8
|
||||||
|
start_score: 10
|
||||||
|
speed_fac: 0.01
|
||||||
|
acc_fac: 0.1
|
||||||
|
die_on_zero: False #True
|
||||||
|
agent_drag: 0.1 # 0.05
|
||||||
|
controll_type: ACC # SPEED
|
||||||
|
aux_reward_max: 1
|
||||||
|
aux_penalty_max: 0.01
|
||||||
|
void_damage: 5 #1
|
||||||
|
agent_draw_path: True
|
||||||
|
---
|
||||||
@@ -0,0 +1,92 @@
|
|||||||
|
name: "DEFAULT"
|
||||||
|
|
||||||
|
# Supported Units:
|
||||||
|
# px: Pixels
|
||||||
|
# em: 1em = Full Width / Height
|
||||||
|
# ct: 100ct = Full Width / Height
|
||||||
|
# rx: pixels relative to width
|
||||||
|
# ry: pixels relative to height
|
||||||
|
# au: 1au = 36px (https://knowyourmeme.com/memes/absolute-unit)
|
||||||
|
#
|
||||||
|
# When no unit is given, we use the folowing defaults
|
||||||
|
# (compatible with legacy behavior)
|
||||||
|
# pos: em
|
||||||
|
# all other: px
|
||||||
|
#
|
||||||
|
# ct is the recommendet unit.
|
||||||
|
# If you need a unit, that is not responsive in regards to width/height, use au / px.
|
||||||
|
|
||||||
|
params:
|
||||||
|
task:
|
||||||
|
task: columbus
|
||||||
|
env_name: ColumbusConfigDefined-v0
|
||||||
|
env_args:
|
||||||
|
observable:
|
||||||
|
- type: State
|
||||||
|
coordsAgent: True
|
||||||
|
speedAgent: True
|
||||||
|
coordsRelativeToAgent: False
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: State
|
||||||
|
coordsAgent: False
|
||||||
|
speedAgent: False
|
||||||
|
coordsRelativeToAgent: True
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: Compass
|
||||||
|
- type: RayCast
|
||||||
|
num_rays: 6
|
||||||
|
chans: [Enemy]
|
||||||
|
entities:
|
||||||
|
- type: RectBarrier
|
||||||
|
num: 1
|
||||||
|
width: 50ct
|
||||||
|
height: 50ct
|
||||||
|
pos: [0ct, 0ct]
|
||||||
|
- type: RectBarrier
|
||||||
|
num: 1
|
||||||
|
width: 50ct
|
||||||
|
height: 50ct
|
||||||
|
pos: [50ct, 50ct]
|
||||||
|
- type: RectBarrier
|
||||||
|
num: 1
|
||||||
|
width: 25rx
|
||||||
|
height: 25ry
|
||||||
|
pos: [0.75em, 30px]
|
||||||
|
- type: RectBarrier
|
||||||
|
num: 1
|
||||||
|
width: 25ry
|
||||||
|
height: 25rx
|
||||||
|
pos: [0.75em, 60px]
|
||||||
|
- type: RectBarrier
|
||||||
|
num: 1
|
||||||
|
width: 20 # defaults to rx (px scaled from x-axis)
|
||||||
|
height: 10 # defaults to ry (px scaled from y-axis)
|
||||||
|
pos: [0.75em, 90px]
|
||||||
|
- type: OnceReward
|
||||||
|
reward: 100
|
||||||
|
radius: 1au
|
||||||
|
pos: [0.3, 0.8] # defaults to em
|
||||||
|
start_pos: [90ct, 20ct]
|
||||||
|
default_collision_elasticity: 0.8
|
||||||
|
start_score: 10
|
||||||
|
speed_fac: 0.01
|
||||||
|
acc_fac: 0.1
|
||||||
|
die_on_zero: False #True
|
||||||
|
agent_drag: 0.1 # 0.05
|
||||||
|
controll_type: ACC # SPEED
|
||||||
|
aux_reward_max: 1
|
||||||
|
aux_penalty_max: 0.01
|
||||||
|
void_damage: 5 #1
|
||||||
|
terminate_on_reward: True
|
||||||
|
agent_draw_path: True
|
||||||
|
clear_path_on_reset: False
|
||||||
|
max_steps: 450 # 1800
|
||||||
|
---
|
||||||
@@ -0,0 +1,108 @@
|
|||||||
|
name: "DEFAULT"
|
||||||
|
|
||||||
|
params:
|
||||||
|
task:
|
||||||
|
task: columbus
|
||||||
|
env_name: Columbus-v1
|
||||||
|
env_args:
|
||||||
|
observable:
|
||||||
|
- type: State
|
||||||
|
coordsAgent: True
|
||||||
|
speedAgent: True
|
||||||
|
coordsRelativeToAgent: False
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: State
|
||||||
|
coordsAgent: False
|
||||||
|
speedAgent: False
|
||||||
|
coordsRelativeToAgent: True
|
||||||
|
coordsRewards: True
|
||||||
|
coordsEnemys: False
|
||||||
|
enemysNoBarriers: True
|
||||||
|
rewardsTimeouts: False
|
||||||
|
include_rand: True
|
||||||
|
- type: Compass
|
||||||
|
- type: RayCast
|
||||||
|
num_rays: 6
|
||||||
|
chans: [Enemy]
|
||||||
|
entities:
|
||||||
|
- type: Ball
|
||||||
|
radius: 16px
|
||||||
|
pos: [0.8, 0.5]
|
||||||
|
speed: [-0.2, -0.1]
|
||||||
|
speed_rand: [0, 0.2]
|
||||||
|
solid: True
|
||||||
|
collision_elasticity: 3
|
||||||
|
elasticity: 1
|
||||||
|
movable: 1
|
||||||
|
collision_changes_speed: True
|
||||||
|
crash_conservation_of_energy: False
|
||||||
|
min_speed: 0.2
|
||||||
|
max_speed: 0.6
|
||||||
|
draw_path: True
|
||||||
|
draw_path_width: 32
|
||||||
|
draw_path_harm: True
|
||||||
|
drag: 0.00001
|
||||||
|
- type: RectGoal # Good
|
||||||
|
height: 1em
|
||||||
|
width: 10ct
|
||||||
|
pos: [97ct, 0ct]
|
||||||
|
skip_agent_col_check: True
|
||||||
|
col: [0, 255, 0]
|
||||||
|
reward: 30
|
||||||
|
solid: True
|
||||||
|
elasticity: 0.6
|
||||||
|
void_collidable: False
|
||||||
|
- type: Goal # Top
|
||||||
|
radius: 7ct
|
||||||
|
pos: [100ct, 0ct]
|
||||||
|
skip_agent_col_check: True
|
||||||
|
col: [0, 255, 0]
|
||||||
|
reward: 30
|
||||||
|
solid: True
|
||||||
|
elasticity: 0.7
|
||||||
|
void_collidable: False
|
||||||
|
- type: Goal # Bottom
|
||||||
|
radius: 7ct
|
||||||
|
pos: [100ct, 100ct]
|
||||||
|
skip_agent_col_check: True
|
||||||
|
col: [0, 255, 0]
|
||||||
|
reward: 30
|
||||||
|
solid: True
|
||||||
|
elasticity: 0.7
|
||||||
|
void_collidable: False
|
||||||
|
- type: RectGoal # Bad
|
||||||
|
height: 1em
|
||||||
|
width: 3ct
|
||||||
|
pos: [0ct, 0ct]
|
||||||
|
skip_agent_col_check: True
|
||||||
|
col: [255, 0, 0]
|
||||||
|
reward: -45
|
||||||
|
solid: True
|
||||||
|
elasticity: 1000
|
||||||
|
void_collidable: False
|
||||||
|
agent_cls: PongAgent
|
||||||
|
agent_attrs:
|
||||||
|
height: 100
|
||||||
|
width: 30
|
||||||
|
movable: False
|
||||||
|
solid: True
|
||||||
|
elasticity: 0.9
|
||||||
|
exception_for_unsupported_collision: False
|
||||||
|
start_pos: [0.05, 0.5]
|
||||||
|
start_score: 0
|
||||||
|
speed_fac: 0.05
|
||||||
|
acc_fac: 0.1
|
||||||
|
die_on_zero: False #True
|
||||||
|
agent_drag: 0
|
||||||
|
controll_type: SPEED
|
||||||
|
aux_reward_max: 0
|
||||||
|
aux_penalty_max: 0
|
||||||
|
void_damage: 0
|
||||||
|
terminate_on_reward: False
|
||||||
|
agent_draw_path: False
|
||||||
|
clear_path_on_reset: False
|
||||||
|
---
|
||||||
@@ -1,4 +1,7 @@
|
|||||||
env_args:
|
name: Example
|
||||||
|
params:
|
||||||
|
task:
|
||||||
|
env_args:
|
||||||
observable:
|
observable:
|
||||||
- type: State
|
- type: State
|
||||||
coordsAgent: True
|
coordsAgent: True
|
||||||
@@ -43,3 +46,4 @@
|
|||||||
aux_reward_max: 1
|
aux_reward_max: 1
|
||||||
aux_penalty_max: 0.1
|
aux_penalty_max: 0.1
|
||||||
void_damage: 5
|
void_damage: 5
|
||||||
|
---
|
||||||
Reference in New Issue
Block a user